Engineering
Lead GPU kernel engineers to optimize CUDA code for faster AI inference
Lead a team of GPU kernel engineers to write and optimize CUDA code for GEMMs, attention mechanisms, and MoE routing. Own technical direction for the kernel roadmap, balancing short-term inference wins with long-term architectural investments. Partner with engineering leadership to align kernel work with Baseten's broader inference stack and ship latency-reducing improvements for AI companies.
Why This Role?
Work on the fastest embeddings solution available and shape the core of Baseten's inference stack
Key Responsibilities
- Lead, grow, and mentor a team of GPU kernel engineers including hiring and career development
- Set technical direction for the kernel roadmap balancing short-term wins with long-term investments
- Partner with Chief Scientist, VP Engineering, and peer engineering leads to align kernel work
- Review CUDA code credibly and unblock team members while staying close to technical details
- Drive model performance optimizations that reduce latency and cost for customer workloads
Requirements
- Hands-on experience writing CUDA kernels for GPU architecture
- Experience optimizing at warp and tensor-core level
- Background in ML systems and production inference
- Ability to lead and mentor elite GPU engineers
- Experience with GEMMs, attention mechanisms, and MoE routing implementations
Required Skills
Keywords
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for an Engineering Manager to lead our GPU Kernel Engineering team, the group responsible for writing the low-level CUDA code that makes Baseten's inference stack faster than anyone else's. This is a player-coach role for someone who has spent years hands-on writing kernels and is now ready to multiply their impact by leading a team of elite GPU engineers. You'll own the technical direction of a team working at the intersection of GPU architecture, ML systems, and production inference. Your engineers write CUDA kernels for GEMMs, attention mechanisms, and MoE routing, optimize at the warp and tensor-core level, and ship improvements that directly reduce latency and cost for the AI companies running their most critical workloads on Baseten. This role is not for someone who wants to step away from the technical work. You'll be close enough to the code to credibly review it, set direction, and unblock your team, while also building the processes, culture, and roadmap that let a world-class kernel team operate at its best. EXAMPLE INITIATIVES Your team owns work like: - Baseten Embeddings Inference: The fastest embeddings solution available https://www.baseten.co/blog/introducing-baseten-embeddings-inference-bei/ - The Baseten Inference Stack https://www.baseten.co/resources/guide/the-baseten-inference-stack/ - Driving model performance optimization https://www.baseten.co/blog/driving-model-performance-optimization-2024-highlights/ RESPONSIBILITIES Team Leadership - Lead, grow, and mentor a team of GPU kernel engineers; own hiring, performance, and career development - Set technical direction for the kernel roadmap, balancing short-term inference wins with long-term architectural investments - Partner closely with the Chief Scientist, VP Engineering, and peer engineering leads to align kernel work with Baseten's broader inference stack strategy - Drive cross-functional collaboration between the kernel team and Model Performance, Capacity, and Infrastructure teams Technical Direction - Establish and maintain a high technical bar for kernel quality, performance, and correctness across the team's output - Review kernel designs and implementations with enough depth to give meaningful feedback on GPU architecture decisions, memory hierarchy tradeoffs, and optimization strategies - Guide the team's approach to profiling and bottleneck identification using tools like Nsight Systems, Nsight Compute, and Torch Profiler - Stay current on the NVIDIA GPU ecosystem (Hopper, Blackwell, and beyond) and translate architectural advancements into team priorities Execution & Culture - Build the processes that allow a highly technical, distributed team to ship with velocity and rigor - Represent the kernel team's work to senior leadership and external audiences including industry conferences - Contribute to Baseten's open-source GPU library presence and technical brand REQUIREMENTS - Proven experience leading a team of GPU or ML systems engineers, with a track record of hiring and developing strong technical talent - Deep personal background in GPU kernel engineering. You have written and shipped production CUDA kernels and can credibly engage with your team's work at a technical level - Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology - Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem - Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership NICE TO HAVE - Hands-on experience with Triton, CUTLASS, or CuTe DSL - Background in LLM inference kernels: attention variants, GEMMs, quantization (FP8/FP4), MoE routing - Open-source contributions to GPU libraries or inference frameworks - Experience presenting technical work at NVIDIA GTC, MLSys, or similar venues BENEFITS - Competitive compensation, including meaningful equity. - 100% coverage of medical, dental, and vision insurance for employee and dependents - Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!) - Paid parental leave - Fertility and family-building stipend through Carrot - Company-facilitated 401(k) - Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities. Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you. At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status. We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
Free account · no credit card · Log in
Pro $9/mo · unlimited applies + AI resume
Frequently asked questions
- Is Engineering at Baseten a remote job?
- This role is based in Remote. See the listing for remote/onsite details.
- What is the salary for Engineering at Baseten?
- The listed pay range for this role is $260k–380k/yr.
- What type of employment is Engineering at Baseten?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Baseten.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Senior Tech Roles Remote Salaries June 2026In-depth analysis of 7 senior tech role remote salaries from Vercel, Airbnb, Stripe, to Notion. Compare with local market and negotiation strategies.
- Sourcing Specialist: The Complete GlobalCurious about becoming a remote Sourcing Specialist? Learn the essential skills, tools, and how to land a USD-paying job—no matter where you live.
- Remote Mobile Developer Jobs July 2026A roundup of USD-paying remote mobile developer jobs from Loker Dollar. Analyze trends and get application tips.
Free account · no credit card · Log in
Pro $9/mo · unlimited applies + AI resume