Langsung ke konten utama
Kembali ke Lowongan

Engineering

Bimbingkan tim GPU Kernel Engineering untuk mengoptimalkan performa AI

Sebagai Engineering Manager di Baseten, Anda akan memimpin tim yang menulis kode CUDA tingkat rendah untuk meningkatkan kecepatan stack inference. Anda akan mengarahkan arah teknis tim, mengoptimalkan performa model, dan mengembangkan proses serta budaya tim. Baseten berfokus pada AI inference untuk perusahaan seperti Cursor dan Notion.

Kenapa Menarik?

Bergabung dengan Baseten untuk memimpin tim GPU kernel engineering dan mengoptimalkan performa AI untuk perusahaan besar.

Tanggung Jawab Utama

  • Membimbing dan mengembangkan tim GPU kernel engineers
  • Mengatur arah teknis untuk roadmap kernel
  • Mengoptimalkan performa model dengan menulis kode CUDA tingkat rendah
  • Mengembangkan proses dan budaya tim untuk operasi terbaik
  • Bekerja sama dengan Chief Scientist dan VP Engineering

Persyaratan

  • Pengalaman dalam menulis kernel CUDA
  • Pengalaman memimpin tim teknis
  • Pemahaman mendalam tentang GPU architecture dan ML systems
  • Kemampuan untuk mengoptimalkan performa model
  • Pengalaman dalam mengembangkan roadmap teknis

Skills Wajib

cudagpu-architectureml-systemsteam-leadershiptechnical-direction

Keywords

engineering-managergpu-kernel-engineeringcudaai-inferenceteam-leadershipremotefull-time
Lihat Deskripsi Asli dari Ashby Job Boards

Deskripsi asli dari Ashby Job Boards

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for an Engineering Manager to lead our GPU Kernel Engineering team, the group responsible for writing the low-level CUDA code that makes Baseten's inference stack faster than anyone else's. This is a player-coach role for someone who has spent years hands-on writing kernels and is now ready to multiply their impact by leading a team of elite GPU engineers. You'll own the technical direction of a team working at the intersection of GPU architecture, ML systems, and production inference. Your engineers write CUDA kernels for GEMMs, attention mechanisms, and MoE routing, optimize at the warp and tensor-core level, and ship improvements that directly reduce latency and cost for the AI companies running their most critical workloads on Baseten. This role is not for someone who wants to step away from the technical work. You'll be close enough to the code to credibly review it, set direction, and unblock your team, while also building the processes, culture, and roadmap that let a world-class kernel team operate at its best. EXAMPLE INITIATIVES Your team owns work like: - Baseten Embeddings Inference: The fastest embeddings solution available https://www.baseten.co/blog/introducing-baseten-embeddings-inference-bei/ - The Baseten Inference Stack https://www.baseten.co/resources/guide/the-baseten-inference-stack/ - Driving model performance optimization https://www.baseten.co/blog/driving-model-performance-optimization-2024-highlights/ RESPONSIBILITIES Team Leadership - Lead, grow, and mentor a team of GPU kernel engineers; own hiring, performance, and career development - Set technical direction for the kernel roadmap, balancing short-term inference wins with long-term architectural investments - Partner closely with the Chief Scientist, VP Engineering, and peer engineering leads to align kernel work with Baseten's broader inference stack strategy - Drive cross-functional collaboration between the kernel team and Model Performance, Capacity, and Infrastructure teams Technical Direction - Establish and maintain a high technical bar for kernel quality, performance, and correctness across the team's output - Review kernel designs and implementations with enough depth to give meaningful feedback on GPU architecture decisions, memory hierarchy tradeoffs, and optimization strategies - Guide the team's approach to profiling and bottleneck identification using tools like Nsight Systems, Nsight Compute, and Torch Profiler - Stay current on the NVIDIA GPU ecosystem (Hopper, Blackwell, and beyond) and translate architectural advancements into team priorities Execution & Culture - Build the processes that allow a highly technical, distributed team to ship with velocity and rigor - Represent the kernel team's work to senior leadership and external audiences including industry conferences - Contribute to Baseten's open-source GPU library presence and technical brand REQUIREMENTS - Proven experience leading a team of GPU or ML systems engineers, with a track record of hiring and developing strong technical talent - Deep personal background in GPU kernel engineering. You have written and shipped production CUDA kernels and can credibly engage with your team's work at a technical level - Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology - Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem - Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership NICE TO HAVE - Hands-on experience with Triton, CUTLASS, or CuTe DSL - Background in LLM inference kernels: attention variants, GEMMs, quantization (FP8/FP4), MoE routing - Open-source contributions to GPU libraries or inference frameworks - Experience presenting technical work at NVIDIA GTC, MLSys, or similar venues BENEFITS - Competitive compensation, including meaningful equity. - 100% coverage of medical, dental, and vision insurance for employee and dependents - Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!) - Paid parental leave - Fertility and family-building stipend through Carrot - Company-facilitated 401(k) - Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities. Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you. At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status. We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Simpan & lacak lamaranmu + dapat match alert

Akun gratis · tanpa kartu kredit · Masuk

Pro Rp39rb/bln · lamar tanpa batas + resume AI

Perusahaan
Baseten
Sumber
Ashby Job Boards
Gaji
$XX,XXX
Lihat selisih gaji remote (USD) vs lokal →
Tipe Pekerjaan
full time
Lokasi
Remote
Kategori
Level
senior
DipostingFresh
9 Jul 2026

Bagikan lowongan ini

Bantu temanmu nemu kerja remote berikutnya.

Pertanyaan yang sering diajukan

Apakah Engineering di Baseten bisa dikerjakan remote?
Posisi ini berlokasi di Remote. Detail remote/onsite ada di deskripsi lowongan.
Berapa gaji untuk Engineering di Baseten?
Rentang gaji yang tercantum untuk posisi ini adalah $260k–380k/yr.
Jenis pekerjaan apa Engineering di Baseten?
Posisi ini adalah pekerjaan full time.
Bagaimana cara melamar?
Klik tombol "Lamar" pada halaman ini untuk menuju halaman aplikasi resmi Baseten.

Jelajahi lebih lanjut

Data & laporan pasar

Riset gaji & permintaan skill dari data lowongan kami sendiri.

Dari blog kami

Simpan & lacak lamaranmu + dapat match alert

Akun gratis · tanpa kartu kredit · Masuk

Pro Rp39rb/bln · lamar tanpa batas + resume AI