AI Engineer
Develop production-ready LLM experiences
Design and implement AI infrastructure, drive LLM experiences, and make data-driven decisions. Work on advanced prompt systems, structured outputs, and complex LLM workflows. Own key AI features from experimentation to live production.
Why This Role?
Take full ownership of key AI features
Key Responsibilities
- Design complex prompt templates with conditional logic
- Implement structured output schemes for predictable AI outputs
- Build evaluation pipelines to score response quality in real-time
- Perform deep debugging of complex LLM chains
- Run systematic experiments across different models
Requirements
- Deep knowledge of Node.js and Next.js
- Proven experience in dynamic prompting
- Experience with LangChain or LlamaIndex
- Familiarity with Langfuse and AI gateways like OpenRouter
Required Skills
Keywords
View Original Description from Landing.jobs
Original description from Landing.jobs
At Ruby Labs (Contractor) Expires at: 2026-08-18 Remote policy: Global remote Ruby Labs is a tech company with a portfolio of consumer products in health, education, and entertainment (100M+ annual users). We'’re seeking a senior AI Engineer (Node.js / Next.js / TypeScript) to shape our AI infrastructure and drive production-ready LLM experiences. You’ll work in a modern stack, making data-driven decisions around model performance, reliability, and cost. You’ll own advanced prompt systems, structured outputs, and complex LLM workflows using LangChain or LlamaIndex. Observability, debugging, and evaluation are core to the role, leveraging Langfuse and AI gateways like OpenRouter to continuously improve model quality and operational efficiency. You’ll take full ownership of key AI features from experimentation to live production. Key Responsibilities: Advanced Prompt Engineering: Designing complex, dynamic prompt templates with conditional logic and efficiently reusing information and context within prompts to maximize generation quality and reasoning. Structured Outputs & Schemas: Implementing various response schemes (JSON mode, function calling, Zod/JSON schemas) to ensure AI outputs are predictable and ready for seamless integration into application logic. Prompt Engineering & Evaluations: Building robust evaluation pipelines and using Langfuse to collect feedback and score the quality of responses in real time. Tracing & Debugging: Performing deep debugging of complex LLM chains using Langfuse traces to identify bottlenecks and optimize for cost, latency, and context window usage. AI A/B Testing: Running systematic experiments across different models via OpenRouter (e.g., comparing Claude 3.5 Sonnet vs. GPT-4o) and analyzing results based on quantitative metrics. Data-Driven Decisions: Making deployment decisions for new prompts or models strictly based on quantitative benchmarks and trace data, rather than intuition. Output Scoring & Analysis: Developing scoring systems to analyze the “Problem → Solution” chain and identify root causes of hallucinations or logic errors using Langfuse analytics. Model Performance & Fine-Tuning: Regularly re-evaluating model performance as new architectures emerge and performing fine-tuning when necessary to meet specific domain requirements. Main requirements Node.js & Next.js: Deep knowledge of the stack to build reliable services and handle complex LLM-generated data. Dynamic Prompting Skills: Proven experience in building prompts where content is highly dependent on input variables and context injection. OpenRouter Experience: Experience working with unified APIs, managing rate limits, and selecting the most cost-effective models for specific tasks. Langfuse (or similar): Understanding of LLM observability principles — setting up tracing, creating test datasets, and integrating scoring systems. Evaluation Methodology: Experience with frameworks like RAGAS or building custom “LLM-as-a-judge” systems. Analytical Mindset: Ability to transform raw generation logs into actionable business metrics and technical insights. Iterative Mindset: Focus on continuous product improvement through constant feedback loops. Fluency in Russian and English. Nice to have Fine-Tuning: Practical experience in fine-tuning models for specific domain tasks or JSON compliance. RAG Architecture: Understanding how to build and optimize Retrieval-Augmented Generation systems, including indexing, retrieval, and re-ranking. Python: Basic knowledge for working with data science scripts or AI evaluation libraries. Benefits & Perks Remote Work Environment: Embrace the freedom to work from anywhere, anytime, promoting a healthy work-life balance. Unlimited PTO: Enjoy unlimited paid time off to recharge and prioritize your well-being, without counting days. Paid National Holidays: Celebrate and relax on national holidays with paid time off to unwind and recharge. Company-provided MacBook: Experience seamless productivity wit
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.