Skip to main content
Back to Jobs

Inference Performance Engineer

Build high-performance inference runtime systems

As an Inference Performance Engineer, you'll optimize throughput, latency, and cost for AI model serving. You'll work on building and improving the inference runtime, designing scheduling and batching systems, and collaborating with hardware teams. Deliverables include optimized inference runtime and improved model serving performance.

Why This Role?

Top-tier compensation and meaningful equity

Key Responsibilities

  • Design and implement scheduling, continuous batching, and KV cache systems
  • Develop low-precision kernels and speculative decoding for improved performance
  • Collaborate with hardware teams on kernel, operator, and graph optimizations
  • Build and maintain benchmarking, profiling, and regression infrastructure

Requirements

  • Software engineering experience with Rust, Go, Python, or C++
  • Understanding of concurrency, memory, and tail latency
  • Experience with model serving frameworks and GPU or ASIC programming
  • Knowledge of modern inference techniques, including transformers and quantization

Required Skills

rustgopythoncudaai inferencesystem optimizationSoftware EngineeringConcurrencyMemory ManagementAIModel Serving

Keywords

Inference PerformanceAI Model ServingConcurrencyGPU ProgrammingASIC ProgrammingModel Serving Frameworks
View Original Description from Ashby Job Boards

Original description from Ashby Job Boards

About the role Serving frontier models at scale requires solving novel systems problems at every layer of the stack. As an Inference Performance Engineer, you'll own the runtime that turns accelerators into a production serving system, optimizing throughput, latency, and cost across thousands of nodes. You'll work alongside hardware and compiler teams operating at the frontier of AI silicon design. What you'll do - Build and improve the inference runtime - Design scheduling, continuous batching, KV cache, and prefill/decode disaggregation - Implement low-precision kernels and speculative decoding - Drive throughput, latency, and cost per token - Collaborate with hardware teams on kernels, operators, and graph optimizations - Own the OpenAI-compatible API surface and serving protocol - Build benchmarking, profiling, and regression infrastructure What you'll need - BS in CS, EE, or related field, or equivalent experience - Software engineering experience: Rust, Go, Python, or C++ - Understanding of concurrency, memory, and tail latency - Understanding of modern inference: transformers, attention, KV cache, batching, speculative decoding, quantization - Experience with model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp, or custom runtimes - GPU or ASIC programming experience: CUDA, ROCm, Triton, or vendor-native toolchains - Experience with low-precision inference (FP8, FP4, INT4) - Profiling and benchmarking experience: Nsight, perf, custom harnesses What we offer - Top-tier compensation structured to recognize and retain the best talent - Meaningful equity - Comprehensive medical, dental, vision, life, and disability insurance - Parental leave for all new parents, including adoptive and surrogate journeys - Flexible PTO - Paid Holidays - Relocation support   Equal Employment Opportunity We're an Equal Opportunity Employer and do not discriminate on the basis of any protected status under applicable law.

Salary Context

Similar Engineering roles on LokerDollar pay around $170k/yr (range $11.194k–999.999k/yr, n=491 active listings).

Hiring at Material Security

Material Security has 4 other active roles on LokerDollar and has been hiring here since Jul 3, 2026 — across Engineering, Data & Analytics.

View all Material Security openings →
Track this application + get a follow-up reminder

Free account · no credit card · Log in

Pro $9/mo · unlimited applies + AI resume

Source
Ashby Job Boards
Salary
Job Type
full time
Location
Remote
Category
Seniority
senior
Posted
May 13, 2026

Share this job

Help a friend find their next remote role.

Frequently asked questions

Is Inference Performance Engineer at Material Security a remote job?
This role is based in Remote. See the listing for remote/onsite details.
What type of employment is Inference Performance Engineer at Material Security?
This is a full time position.
How do I apply?
Click the "Apply" button on this page to go to the official application at Material Security.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.

From the blog

Track this application + get a follow-up reminder

Free account · no credit card · Log in

Pro $9/mo · unlimited applies + AI resume