Duck Duck Go

Director, Site Reliability Engineering

Remote (Remote) remote Mid Salary not listed
remote Mid level Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

Who We Are
Hi, we're DuckDuckGo, the online protection company and remote-first team of 300+ on a mission to raise the standard of trust online. Founded in 2008 and profitable since 2014, annual revenue now exceeds $100m USD and millions use our browser on on Mac, Windows, iOS, and Android, our search engine, and the DuckDuckGo subscription. We also offer private, useful, and optional AI, including Duck.ai, which lets you chat privately with ChatGPT, Claude, and other AIs, all in one place. Our culture of trust, inclusivity, and empowered project management underpins everything we do, where each team member takes full ownership of their projects, from scoping and execution to postmortem. If you're seeking end-to-end ownership of your work, you've come to the right place!

Your Team and Role
Working on the Site Reliability Team, you'll help build and maintain world-class infrastructure to meet the needs of millions of users protecting their privacy online. You'll utilize high-level languages like Perl, Go, TypeScript, or Python and work on related projects. Recent projects include:
Ensuring our Duck.ai product meets our reliability standards and minimizing user friction on failures

Scaling up our own index infrastructure to handle billions of documents

Create anti fraud verifications that respect users privacy

 
As Director, Site Reliability Engineering, you'll dive deep into complex operational challenges, including software, systems, automation, and process analysis. We are looking for candidates who can read, write, troubleshoot, and deploy all types of software to help us tackle the reliability challenges of large-scale deployments.
 
About You
10+ years relevant professional experience in reliability, platform, infrastructure, or software engineering, including 4+ years leading SRE teams.

Experience participating in a 24x7 on-call rotation for a large-scale deployment.

Ability to lead and collaborate on high-impact and complex projects from proposal through postmortem.

Proficient in AI-driven development, including designing and implementing agentic workflows

Skills to wrangle vague problems, propose innovative solutions, and execute them with a strong focus on metrics.

Experience developing effective tools, services, alerts, and responses to identify and address reliability risks.

Investigative ability to root-cause sources of instability in high-traffic, distributed systems.

Deep experience administering and troubleshooting Linux and web technologies.

Ability to implement automation around infrastructure provisioning and configuration management to prioritize efficiency, scalability, and reliability.

Foresight to help identify the future technical direction of our deployment with the goal of improving reliability and performance.

Advanced programming skills enabling close partnership with software engineers to triage production issues and identify appropriate remediation, including code changes and performance considerations.

Ability to leverage cloud-native services and architectures to enhance reliability and scalability, with hands-on experience packaging and deploying applications using Docker and Docker Compose.

Compensation
$243,800 USD annually and stock options. Compensation is transparent across the organization, and all team members within the same professional level and global region receive the same compensation.
 
Eligibility for company-sponsored health benefits is limited to team members based in the United States. This program does not extend to team members located in other countries, such as Canada or the UK.
 
Our Team Member Support Guide explains how we prioritize your wellbeing including paid parental leave, office setup, and co-working allowances.
 
Hiring Process
Hiring works best when it's a two-way street. Learn how we help you get to know DuckDuckGo, envision your future role here, and find out more about how we hire.
 
Diversity, Equity and Inclusion
DuckDuckGo provides equal work opportunities to all team members and applicants, and it prohibits discrimination and harassment of any type on the basis of race, color, ethnicity, caste, religion, age, sex (including pregnancy), national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by our policies or federal, state, or local laws.
 
We want to ensure that our hiring process is accessible. If you need reasonable accommodation for any part of the application process because of a medical condition or disability, please send an email to

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Mid
Technology & IT Javascript Project Management Python Remote Collaboration Mid level remote

Likely questions

  1. Tell us about work you have done that is close to the Director, Site Reliability Engineering role.
  2. How would you approach your first 30 days at Duck Duck Go?
  3. Which of Javascript, Project Management and Python have you used recently, and what did it help you achieve?
  4. Describe a time you solved a problem without waiting to be told exactly what to do.
  5. How do you stay organised and communicate clearly when working remotely?

Prepare before the call

  • A recent example that proves your experience with Javascript, Project Management and Python.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Director, Site Reliability Engineering role because I can bring practical experience in Javascript, Project Management and Python, learn the team quickly, and contribute to the outcomes Duck Duck Go needs from this hire.

Related jobs.

More roles from this company or category.