Runware

Senior Site Reliability Engineer

France, Germany, Italy full-time Senior Salary not listed
full-time Senior level Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.
As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.
What you’ll do

Own and improve the reliability, availability and performance of critical production services across the Runware platform

Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards

Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation

Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements

Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience

Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

Requirements

Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role

Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure

Have experience designing and operating observability systems using metrics, logs and distributed tracing

Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil

Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP

Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation

Bonus

Experience operating high-throughput or low-latency APIs and distributed systems

Experience with bare-metal infrastructure, GPU environments or AI and ML workloads

Experience with RabbitMQ or other distributed messaging and queueing systems

Experience operating MySQL, Redis, ClickHouse or similar production data systems

Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments

Experience building automated scaling, capacity management or self-healing systems

Benefits
We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.
Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

Generous paid time off – vacation, sick days, public holidays

Meaningful stock options – share in the upside you create

Remote-first setup – work from home anywhere we can employ you

Flexible hours – own your schedule outside core collaboration blocks

Family leave – paid maternity, paternity, and caregiver time

Company retreats – twice-yearly gatherings in inspiring locations

Originally posted on Himalayas

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Senior
Technology & IT MySQL Operations PHP Python Remote Collaboration Senior level

Likely questions

  1. Tell us about work you have done that is close to the Senior Site Reliability Engineer role.
  2. How would you approach your first 30 days at Runware?
  3. Which of MySQL, Operations and PHP have you used recently, and what did it help you achieve?
  4. How have you led people, improved a process, or made a hard decision in a previous role?
  5. How do you handle busy days, changing priorities, or pressure at work?

Prepare before the call

  • A recent example that proves your experience with MySQL, Operations and PHP.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Senior Site Reliability Engineer role because I can bring practical experience in MySQL, Operations and PHP, learn the team quickly, and contribute to the outcomes Runware needs from this hire.

Related jobs.

More roles from this company or category.