Site Reliability Engineer
Develop system observability using metrics logs events and traces
As a Site Reliability Engineer at Trainline, you will develop an understanding of system architecture dependencies and failure modes across the platform. You will participate in production incident response supporting investigation mitigation communication and coordinated service restoration. You will contribute to post-incident reviews and follow-up actions to improve reliability scalability and resilience while taking part in the SRE on-call...
Why This Role?
Work on a platform serving over 135 million monthly visits with £6.3 billion in annual ticket sales
Key Responsibilities
- Develop understanding of system architecture dependencies and failure modes across the Trainline platform
- Participate in production incident response supporting investigation mitigation communication and coordinated service restoration
- Contribute to post-incident reviews and follow-up actions to improve reliability scalability and resilience
- Take part in the SRE on-call rotation
- Design build and maintain observability using metrics logs events and traces
Requirements
- Solid production experience
- Growth mindset
- Willingness to challenge and be challenged
- Experience with AWS cloud-native architecture
- Familiarity with modern CI/CD pipelines and DevOps practices
Required Skills
Keywords
View Original Description from Ashby Job Boards
Original description from Ashby Job Boards
About us We are champions of rail, inspired to build a greener, more sustainable https://www.thetrainline.com/terms/sustainability-faqs future of travel. Trainline enables millions of travellers to find and book the best value tickets across carriers, fares, and journey options through our highly rated mobile app, website, and B2B partner channels. Great journeys start with Trainline 🚄 Now Europe’s number 1 downloaded rail app, with over 135 million monthly visits and £6.3 billion in annual ticket sales, we collaborate with 270+ rail and coach companies in over 40 countries. We want to create a world where travel is as simple, seamless, eco-friendly and affordable as it should be. Today, we're a FTSE 250 company driven by our incredible team of over 1,000 Trainliners from 50+ nationalities, based across London, Paris, Barcelona, Milan, Edinburgh and Madrid. With our focus on growth in the UK and Europe, now is the perfect time to join us on this high-speed journey. Introducing Reliability & Operations Engineering 👋 Trainline is a fast-growing tech company powering world-class digital journeys for millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices. The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery, respond to incidents, and continuously strengthen system reliability. We’re looking for a mid-level Site Reliability Engineer to help drive this forward. You’ll bring solid production experience, a growth mindset, and a willingness to challenge and be challenged — contributing to platform reliability while developing broader technical ownership with support from senior engineers. As an SRE at Trainline, you'll be working on...🚄 - Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform - Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration - Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience - Taking part in the SRE on-call rotation - Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis - Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD) - Ensuring relevant operational data is surfaced quickly and clearly during live incidents - Making informed tooling and technology choices using SRE principles, balancing team and business needs - Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling - Collaborating with product engineering teams to ensure services are operationally ready and deployed safely - Advising on reliability and resilience practices - Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals - Prioritising work effectively and collaborating using agile processes to deliver against team and business goals Our Tech Stack 🔑 - AWS - New Relic - ELK stack - Grafana - Incident.io - Docker, ECS - Terraform - Github Actions We'd love to hear from you if you have...🔍 - Experience of SRE concepts such as SLI, SLO and error budgets. - Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar - Experience working with cloud providers (preferably AWS). - Experience troubleshooting Linux operating systems. - Experience of scripting in at least one language (preferably Python) - Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts. - Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling). - Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions. - Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform. More information: Enjoy fantastic perks like private healthcare & dental insurance, a generous work from abroad policy, 2-for-1 share purchase plans, an EV Scheme to further reduce carbon emissions, extra festive time off, and excellent family-friendly benefits. We prioritise career growth with clear career paths, transparent pay bands, personal learning budgets, and regular learning days. Jump on board and supercharge your career from day one! We're operating a hybrid model and ask that Trainliners work from the office a minimum of 60% of their time over a 12-week period. We also have a 28-day Work from Abroad policy. Our values represent the things that matter most to us and what we live and breathe everyday, in everything we do: - 💭 Think Big - We're building the future of rail - ✔️ Own It - We focus on every customer, partner and journey - 🤝 Travel Together - We're one team - ♻️ Do Good - We make a positive impact We know that having a diverse team makes us better and helps us succeed. And we mean all forms of diversity - gender, ethnicity, sexuality, disability, nationality and diversity of thought. That's why we're committed to creating inclusive places to work, where everyone belongs and differences are valued and celebrated. Interested in finding out more about what it's like to work at Trainline? Why not check us out on LinkedIn https://www.linkedin.com/company/trainline/, Instagram https://www.instagram.com/lifeattrainline/ and Glassdoor https://www.glassdoor.co.uk/Overview/Working-at-Trainline-EI_IE249203.11,20.htm!
Free account · no credit card · Log in
Pro $9/mo · unlimited applies + AI resume
Frequently asked questions
- Is Site Reliability Engineer at Trainline a remote job?
- This role is based in Remote. See the listing for remote/onsite details.
- What is the salary for Site Reliability Engineer at Trainline?
- The listed pay range for this role is £55k–63k/yr.
- What type of employment is Site Reliability Engineer at Trainline?
- This is a full time position.
- How do I apply?
- Click the "Apply" button on this page to go to the official application at Trainline.
Explore related
Market data & reports
Salary & skill-demand research built from our own listings data.
- Indonesia IT Jobs vs Global Remote (2026)Primary analysis of 2,049 listings: methodology, classification rules, downloadable datasets.
- AI-Skill Demand: Indonesia vs Global Remote (2026)10,000+ postings, taxonomy-first classifier, Wilson CIs, pre-registered before analysis.
- Indonesia Hiring Report: Tech vs Non-TechJob demand by field from aggregate open-job counts — never individual listings.
- Indonesia Salary BenchmarkAggregate salary ranges across roles, with open methodology and dataset.
- Indonesia Quarterly Labor Market ReportLayoffs, funding, salaries & skills per quarter — open aggregates.
- Remote Market Reports by RoleAuto-generated per role family — skills, seniority, companies, salary.
- Global Remote Salary BenchmarkAnnual salary by role & currency, plus the share of listings open worldwide.
From the blog
- Senior Tech Roles Remote Salaries June 2026In-depth analysis of 7 senior tech role remote salaries from Vercel, Airbnb, Stripe, to Notion. Compare with local market and negotiation strategies.
- Sourcing Specialist: The Complete GlobalCurious about becoming a remote Sourcing Specialist? Learn the essential skills, tools, and how to land a USD-paying job—no matter where you live.
- Remote Mobile Developer Jobs July 2026A roundup of USD-paying remote mobile developer jobs from Loker Dollar. Analyze trends and get application tips.
Free account · no credit card · Log in
Pro $9/mo · unlimited applies + AI resume
