Research Engineer – Evals
Build evaluation systems to measure Firecrawl's data extraction quality
You will design metrics, build pipelines, and generate datasets to evaluate whether Firecrawl's outputs are accurate across diverse websites and formats. This includes creating benchmarks that reflect real-world usage, integrating evaluations into CI/CD, and owning the feedback loop from output quality to model and product decisions. You will work on eval infrastructure from scratch, covering scrape, crawl, extract, and map functions.
Why This Role?
Own the eval stack from scratch and directly impact product and model decisions
Key Responsibilities
- Design metrics to evaluate Firecrawl's output quality across scrape, crawl, extract, and map functions
- Build evaluation pipelines and curate datasets for benchmarking
- Integrate evals into CI/CD to catch regressions before shipping
- Design benchmark datasets reflecting real-world website diversity including edge cases
- Own the feedback loop from output quality to model and product decisions
Requirements
- 3+ years in ML engineering, applied AI, or data quality with production systems
- Experience building evaluation systems or benchmarks for AI/ML outputs
- Ability to design metrics and curate datasets for LLM-ready data assessment
- Experience with pipelines and CI/CD integration for testing
- Background in measuring data quality in complex, variable environments
Required Skills
Indonesia Context
- Working Hours Overlap:
- Minimal overlap — opposite hours
Keywords
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
RESEARCH ENGINEER — EVALS You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our core promise — convert any URL into clean, structured, LLM-ready data reliably — is hard to measure rigorously across millions of different websites, formats, and edge cases. As we layer in models and agent workflows, the question "did that work?" gets harder, not easier. This isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what "good" actually means and have the engineering depth to measure it, this is the role. Salary Range: $160,000 to $240,000/year (Range shown is for U.S.-based employees in San Francisco, CA. Compensation outside the U.S. is adjusted fairly based on your country's cost of living.) Equity Range: Up to 0.10% Location: San Francisco, CA or Remote (Americas, UTC-3 to UTC-10) Job Type: Full-Time Experience: 3+ years in ML engineering, applied AI, or data quality — with production systems Visa: US Citizenship/Visa required for SF; N/A for Remote ABOUT FIRECRAWL Firecrawl is the easiest way to extract data from the web. Developers use us to reliably convert URLs into LLM-ready markdown or structured data with a single API call. In just a year, we've hit 8 figures in ARR and 120k+ GitHub stars by building the fastest way for developers to get LLM-ready data. We're a small, fast-moving, technical team building essential infrastructure superintelligence will use to gather data on the web. We ship fast and deep. WHAT YOU’LL DO Build the eval stack from scratch. Design and own the systems that measure whether Firecrawl's outputs are actually good — across scrape, crawl, extract, and map. That means defining metrics, building pipelines, curating datasets, and integrating evals into CI/CD so regressions get caught before they ship. You build the infra yourself because you're the one who needs it to work. Design benchmarks that reflect reality. Our outputs need to hold up across millions of websites — SPAs, paywalled content, dynamic rendering, structured and unstructured formats. You'll build benchmark datasets that cover the real distribution of what our customers send us, including the edge cases that break naive approaches. Ground truth doesn't come for free — you'll design the collection and labeling systems too. Own LLM-as-judge pipelines. You'll design and validate automated judges that score extraction quality at scale, know the failure modes of LLM-based evaluation, and build the human review tooling needed when automation isn't enough. You understand the difference between an eval that measures something real and one that just flatters the system. Close the loop with models and RL. Evals here aren't a reporting layer — they're a training signal. You'll work closely with the RL and Search/IR research engineers to turn quality measurements into reward signals and feedback loops that make models meaningfully better. Your benchmarks directly influence what gets trained next. Run fast experiments and communicate clearly. You design experiments that test meaningful hypotheses, run them quickly, and make decisions based on results. When you have findings, anyone on the team can understand what they mean — no decoder ring required. WHAT WE'RE LOOKING FOR Builds their own eval infrastructure. You don't wait for tooling to appear. You write the pipelines, curate the datasets, design the rubrics, and validate the judges yourself — because you understand that infra choices directly affect what you're actually measuring. You've run evals at scale and debugged the places where they lie. Knows what "good" means for unstructured web data. You've worked with messy, real-world data before. You understand why markdown quality is hard to define, why structured extraction fidelity varies by schema, and why naive string-match metrics miss the point. You have strong opinions about what a useful benchmark actually looks like — and the rigor to validate them. Fluent in LLM evaluation methodology. You understand LLM-as-judge systems, their correlation with human judgment, and where they break down. You've designed rubrics that hold up under adversarial inputs, built human review pipelines that scale, and know how to measure inter-rater agreement. You're not fooled by evals that only look good in aggregate. Production-minded. You care about whether your evals reflect real production behavior, not just offline benchmarks. You've worked on systems serving real traffic and made hard tradeoffs between evaluation depth, coverage, and cost. A benchmark that doesn't represent what customers actually send isn't a benchmark worth maintaining. Fast and clear. You'd rather run three rough experiments this week than one polished one next month. When you have results, anyone on the team can understand what they mean — and what to do next. Backgrounds that tend to do well: ML engineers who've built eval or data quality systems at AI labs or applied teams. Engineers who've worked on LLM fine-tuning or RLHF pipelines and understand how feedback quality drives model improvement. People who've worked at the intersection of data infrastructure and model development. Anyone who's been the person on the team asking "but how do we know this actually works?" WHAT WE'RE NOT LOOKING FOR Benchmark runners. If your eval experience is running existing frameworks on existing benchmarks and reporting numbers, this isn't the right fit. We need someone who builds the frameworks and defines the benchmarks. People who treat evals as an afterthought. If your default workflow is to build first and evaluate later — or to treat pass rates as a proxy for actual quality — you'll struggle here. Evals are a first-class product, not a QA gate. Researchers who need a platform team. If you expect pipelines, datasets, and labeling infrastructure to exist before you can be productive, you'll be frustrated. You build the tools you need. Slow iterators. If your standard experiment cycle is measured in weeks, not days, you'll struggle with the pace. We need someone who can design, run, and interpret a meaningful experiment within a day or two. BONUS POINTS - Any other niche expertise and skills - Previous experience at a scraping, automation, or security-focused startup - Ex-founder BENEFITS & PERKS AVAILABLE TO ALL EMPLOYEES - Salary that makes sense — $160,000-240,000/year (U.S.-based), based on impact, not tenure - Own a piece — Up to 0.1% equity in what you're helping build - Unlimited PTO — Minimum 3 weeks off encouraged; take the time you need to recharge - Parental leave — 12 weeks fully paid, for moms and dads - Wellness stipend — $100/month for the gym, therapy, massages, or whatever keeps you human - Learning & Development - Expense up to $150/year toward anything that helps you grow professionally - Team offsites — A change of scenery, minus the trust falls - Sabbatical — 3 paid months off after 4 years, do something fun and new AVAILABLE TO US-BASED FULL-TIME EMPLOYEES - Full coverage, no red tape — Medical, dental, and vision (100% for employees, 50% for spouse/kids) — no weird loopholes, just care that works - Life & Disability insurance — Employer-paid short-term disability, long-term disability, and life insurance — coverage for life's curveballs - Supplemental options — Optional accident, critical illness, hospital indemnity, and voluntary life insurance for extra peace of mind - Doctegrity telehealth — Talk to a doctor from your couch - 401(k) plan — Retirement might be a ways off, but future-you will thank you - Pre-tax benefits — Access to FSAs and commuter benefits to help your wallet out a bit - Pet insurance — Because fur babies are family too AVAILABLE TO SF-BASED EMPLOYEES - SF HQ perks — Snacks, drinks, team lunches, and the occasional burst of chaotic startup energy INTERVIEW PROCESS 1. Application Review – Send us your stuff, and a quick note on why you're excited 2. Intro Chat (~25 min) – Quick alignment call with a member of our team 3. Technical Interview (~1 hr) – Tackle a small challenge 4. Interview with Founders (~30 min) – Culture, vision, and long-term fit 5. Paid Work Trial (1–2 weeks) – Work on something real with us 6. Decision – We move fast If you’ve ever wanted to own a product-critical system and build alongside founders, this is your moment. Apply now and let’s talk.
Salary Context
Similar Engineering roles on LokerDollar pay around $170k/yr (range $11.194k–999.999k/yr, n=491 active listings).
Free account · no credit card · Log in
Pro $9/mo · unlimited applies + AI resume
Hiring in US only
This employer appears to hire only in the region above. Confirm you're eligible to be hired there before applying.
Frequently asked questions
- Is Research Engineer – Evals at Firecrawl a remote job?
- Yes. Research Engineer – Evals at Firecrawl is a fully remote role open to candidates worldwide.
- What is the salary for Research Engineer – Evals at Firecrawl?
- The listed pay range for this role is $210k–275k/yr.
- What type of employment is Research Engineer – Evals at Firecrawl?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Firecrawl.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- 10 Remote Jobs Nobody Wants (Yet!)Are there remote openings with few applicants? We'll show you 10 high-demand, low-supply positions – and why your skills might be a perfect fit.
- Remote USD Jobs in July 2026: Top OpeningsDiscover the latest remote USD job openings for July 2026, featuring top companies like Tailor, Vanta, and Kyndryl, with salaries up to $166,000 per year.
- Remote Freelance Jobs June 2026: Trends & PayA global look at seven new remote freelance gigs, pay ranges, a contrarian take, and practical tips for landing USD‑paid work.
Free account · no credit card · Log in
Pro $9/mo · unlimited applies + AI resume
