Data Engineer
Maintain and optimize large-scale data pipelines for web scraping infrastructure
As a Data Engineer at Wynd Labs, you will maintain, optimize, and improve large-scale data pipelines and infrastructure supporting web data collection for AI training. Your work will focus on scalability, reliability, and performance across data collection, processing, transformation, validation, and delivery. You'll operate within a lean, fast-moving team building infrastructure for massive public web data access.
Why This Role?
Work on infrastructure powering dataset creation for frontier AI labs with real impact on open web data
Key Responsibilities
- Maintain and optimize large-scale data pipelines for web scraping and processing
- Work with distributed systems and task queues like Celery, Kafka, or RabbitMQ
- Manage containerized workloads using Docker and Kubernetes with Helm charts
- Ensure reliable deployment and autoscaling of scraping and processing jobs on Linux servers
- Implement CI/CD workflows using GitHub Actions or ArgoCD for data pipelines
- Write scalable APIs and handle complex analytical queries in columnar warehouses
Requirements
- Advanced Python with strong async programming and multiprocessing skills
- Hands-on experience with high-volume web scraping using proxies and anti-bot evasion
- Experience designing and operating distributed data pipelines with task queues
- Practical experience with columnar data warehouses such as Databend, ClickHouse, or BigQuery
- Proficiency in Docker, Kubernetes, and managing deployments with Helm
- Comfort with Linux server management, debugging performance issues, and bare-metal operations
Required Skills
Indonesia Context
- Working Hours Overlap:
- Minimal overlap — opposite hours
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
Who We Are: We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs. We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI. The Role: We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance. This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads. Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the team. Who You Are: - Bachelor’s degree or equivalent work experience - Python (advanced) — strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs - Web scraping at scale — hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media) - Distributed data pipelines — experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar) - Data warehousing — practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses - Docker & Kubernetes — containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads - Linux & bare-metal ops — comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions - CI/CD for data workflows (GitHub Actions, ArgoCD) - Writing Scalable API What You'll Be Doing: - Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability. - Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets. - Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements. - Monitor and troubleshoot data pipeline issues, identify data quality concerns, and help implement timely fixes to maintain data accuracy and operational continuity. - Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented. - Participate in research and development projects to improve the Company’s data products and workflows. Why Work With Us: - Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people. - Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better. We prioritize low ego and high output. This is a fully remote team. - Compensation. You’ll receive a competitive salary, benefits and equity package.
Salary Context
Similar Engineering roles on LokerDollar pay around $170k/yr (range $855–1000k/yr, n=887 active listings).
Openness not stated by employer — check the listing
The listing does mention EST work hours — check whether that overlap works for you.
Frequently asked questions
- Is Data Engineer at Wynd Labs a remote job?
- Yes. Data Engineer at Wynd Labs is a fully remote role open to candidates worldwide.
- What type of employment is Data Engineer at Wynd Labs?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Wynd Labs.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Remote ≠ Remote: The Skills That Open Global Work to Indonesians (2026)12,891 remote listings: the highest-paid coding skills are the most geo-locked for Indonesia-based applicants. CC BY 4.0 aggregate dataset.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesian Remote Work Salary & Demand IndexHow much of the global remote job corpus is open to Indonesia, and what it pays (USD) by role.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Remote Jobs with USD Salaries: WhatExplore remote job opportunities with USD salaries and discover what to expect in terms of pay, benefits, and requirements for various roles.
- The Shift in Remote Hiring: What CIOs AreAnalyzing the latest remote hiring trends and CIO interview strategies to help you land your next USD-paying role in a competitive global market.
- Are You AI-Ready? Remote Jobs and Global PayIs the AI era passing you by? Explore the latest global remote job openings and learn how to benchmark your salary against international standards.