Langsung ke konten utama
Kembali ke Lowongan

Inference Performance Engineer

Optimalkan kinerja model AI di skala besar untuk Material Security

Anda akan mengembangkan dan memantau runtime inference yang mengubah akselerator menjadi sistem produksi. Anda akan bekerja sama dengan tim hardware dan compiler untuk meningkatkan throughput, latency, dan biaya di ribuan node. Anda akan mengoptimalkan berbagai komponen seperti scheduling, batching, dan low-precision kernels.

Kenapa Menarik?

Dapatkan kompensasi terbaik dan equity yang signifikan di Material Security

Tanggung Jawab Utama

  • Bangun dan perbaiki runtime inference
  • Rancang scheduling, batching, dan KV cache
  • Implementasikan low-precision kernels dan speculative decoding
  • Optimalkan throughput, latency, dan biaya per token
  • Kerjasama dengan tim hardware untuk optimasi kernel dan operator
  • Pemilik API dan protokol serving yang kompatibel dengan OpenAI

Persyaratan

  • Pendidikan S1 di bidang CS, EE, atau setara
  • Pengalaman software engineering dengan Rust, Go, Python, atau C++
  • Paham tentang concurrency, memory, dan tail latency
  • Paham tentang modern inference: transformers, attention, KV cache, batching
  • Pengalaman dengan model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM
  • Pengalaman programming GPU atau ASIC: CUDA, ROCm, Triton

Skills Wajib

rustgopythoncudaai inferencesystem optimization

Keywords

inference performanceai model servingrustgopythoncudaremotefull-time
Lihat Deskripsi Asli dari Ashby Job Boards

Deskripsi asli dari Ashby Job Boards

About the role Serving frontier models at scale requires solving novel systems problems at every layer of the stack. As an Inference Performance Engineer, you'll own the runtime that turns accelerators into a production serving system, optimizing throughput, latency, and cost across thousands of nodes. You'll work alongside hardware and compiler teams operating at the frontier of AI silicon design. What you'll do - Build and improve the inference runtime - Design scheduling, continuous batching, KV cache, and prefill/decode disaggregation - Implement low-precision kernels and speculative decoding - Drive throughput, latency, and cost per token - Collaborate with hardware teams on kernels, operators, and graph optimizations - Own the OpenAI-compatible API surface and serving protocol - Build benchmarking, profiling, and regression infrastructure What you'll need - BS in CS, EE, or related field, or equivalent experience - Software engineering experience: Rust, Go, Python, or C++ - Understanding of concurrency, memory, and tail latency - Understanding of modern inference: transformers, attention, KV cache, batching, speculative decoding, quantization - Experience with model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp, or custom runtimes - GPU or ASIC programming experience: CUDA, ROCm, Triton, or vendor-native toolchains - Experience with low-precision inference (FP8, FP4, INT4) - Profiling and benchmarking experience: Nsight, perf, custom harnesses What we offer - Top-tier compensation structured to recognize and retain the best talent - Meaningful equity - Comprehensive medical, dental, vision, life, and disability insurance - Parental leave for all new parents, including adoptive and surrogate journeys - Flexible PTO - Paid Holidays - Relocation support   Equal Employment Opportunity We're an Equal Opportunity Employer and do not discriminate on the basis of any protected status under applicable law.

Konteks Gaji

Posisi Engineering serupa di LokerDollar dibayar sekitar $170k/yr (kisaran $11.194k–999.999k/yr, dari 491 listing aktif).

Perekrutan di Material Security

Material Security punya 4 lowongan aktif lain di LokerDollar dan telah merekrut di sini sejak 3 Jul 2026 — di kategori Engineering, Data & Analytics.

Lihat semua lowongan Material Security →
Simpan & lacak lamaranmu + dapat pengingat follow-up

Akun gratis · tanpa kartu kredit · Masuk

Pro Rp39rb/bln · lamar tanpa batas + resume AI

Sumber
Ashby Job Boards
Gaji
Tipe Pekerjaan
full time
Lokasi
Remote
Kategori
Level
mid
Diposting
13 Mei 2026

Bagikan lowongan ini

Bantu temanmu nemu kerja remote berikutnya.

Pertanyaan yang sering diajukan

Apakah Inference Performance Engineer di Material Security bisa dikerjakan remote?
Posisi ini berlokasi di Remote. Detail remote/onsite ada di deskripsi lowongan.
Jenis pekerjaan apa Inference Performance Engineer di Material Security?
Posisi ini adalah pekerjaan full time.
Bagaimana cara melamar?
Klik tombol "Lamar" pada halaman ini untuk menuju halaman aplikasi resmi Material Security.

Jelajahi lebih lanjut

Data & laporan pasar

Riset gaji & permintaan skill dari data lowongan kami sendiri.

Dari blog kami

Simpan & lacak lamaranmu + dapat pengingat follow-up

Akun gratis · tanpa kartu kredit · Masuk

Pro Rp39rb/bln · lamar tanpa batas + resume AI