Supersourcing

SRE (Site Reliability Engineer)

Remote, United States remote Entry Salary not listed
remote Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

About the job
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that the servicesboth our internally critical and our externally-visible systemshave reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally, SREs will keep an ever-watchful eye on our systems capacity and performance.
As a Site Reliability Engineer, you will have the opportunity to manage the complex challenges of scale which are unique to Digitization, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. You will provide scalable, reliable, durable, and secure applications for our customers and internal users. You will help build highly reliable applications using a customer-first approach while innovating technically. You will understand our customer's needs and how we can meet them.
Responsibilities
Work with the Site Reliability Engineering team, Development team, and other partner teams to ensure that applications reliability, efficiency, and performance meets our customer's needs, while keeping the service's operation's reliable, scalable, and automated.

Develop and implement projects that improve system reliability, efficiency, and performance

Partner with development teams on feature launches to ensure our customers are delivered reliable and scalable functionality.

Build a deep knowledge on production infrastructure and using that to debug distributed systems problems and identify improvements to the system.

Operations, SLO, SLA management

Metrics reporting and progress tracking

Be on-call, responding to and managing incidents.

Observability (Alarms, monitoring, synthetics).

Error management

Qualifications
Bachelor's degree in Computer Science or a related engineering degree

8+ years of IT industry experience

Strong Experience in

Java, Springboot, Nodejs, microservices, RDBMS, NoSQL

AWS EC2, S3, Lambda, IAM, ECS, EKS, SQS, Kinesis

Observability using Splunk, NewRelic

Infrastructure as Code using terraform

APIs and event-driven approaches

Security patterns

Unix/Linux systems administration. Familiar with Docker is a must.

Strong Experience in analysing and troubleshooting large-scale distributed systems. Quick reaction on high severity customer impacts.

Ability to debug and optimize code and automate routine tasks

Knowledge in modern software engineering practices and tools - Agile and DevOps

Strong communication skill and the ability to explain complex technical matters in an easy-to-understand way.

Originally posted on Himalayas

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Role
Technology & IT Data Analysis Operations Project Management Sre Site remote

Likely questions

  1. Tell us about work you have done that is close to the SRE (Site Reliability Engineer) role.
  2. How would you approach your first 30 days at Supersourcing?
  3. Which of Data Analysis, Operations and Project Management have you used recently, and what did it help you achieve?
  4. Describe a time you solved a problem without waiting to be told exactly what to do.
  5. How do you stay organised and communicate clearly when working remotely?

Prepare before the call

  • A recent example that proves your experience with Data Analysis, Operations and Project Management.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the SRE (Site Reliability Engineer) role because I can bring practical experience in Data Analysis, Operations and Project Management, learn the team quickly, and contribute to the outcomes Supersourcing needs from this hire.

Related jobs.

More roles from this company or category.