Site Reliability Engineer

TikTok
San Jose, CA
Job Description
Role Overview

We are seeking Site Reliability Engineers to join our team in San Jose, California. As a Site Reliability Engineer, you will provide site reliability engineering support to ensure the highest level of availability of large-scale, fault-tolerant systems. You will assist the team in improving the reliability, scalability, and release cycle of infrastructure services from inception, design, and development.

What You Will Do

Your day-to-day responsibilities will include measuring and monitoring availability, latency, and overall service health, practicing sustainable user support, incident response, and blameless postmortems, building tools, automations, visualizations, and monitors to improve the reliability and scalability of services, and sharing on-call responsibility and troubleshooting problems across a wide array of services and functional areas.

Why It Might Be a Fit

We are looking for talents who are willing to learn and put in the effort to solve problems. Our challenges are not your regular day-to-day problems - you’ll be part of a team that’s developing new solutions to new challenges. It’s working fast, at scale, and we’re making a difference.

Requirements

  • Master's degree or foreign equivalent degree in Computer Science, Engineering (any), Information Systems, Mathematics, or a related field, and 1 year of related work experience
  • Bachelor's degree or foreign equivalent degree in Computer Science, Engineering (any), Information Systems, Mathematics, or a related field, and 3 years of related work experience
  • 1 year of experience in providing functionality and reliability support for critical site components by measuring and monitoring availability, latency, and overall system health
  • 1 year of experience in monitoring system activity and resolving system issues
  • 1 year of experience in coordinating and monitoring data services operations, including SLA management and system deployment
  • 1 year of experience in analyzing error logs to identify issues and working with service owners to resolve issues, document their origins and develop future prevention mechanisms
  • 1 year of experience in creating and maintaining clear runbook instructions for services to use for alerts, troubleshooting and resolution

Benefits

  • Salary Range: $226138 - $316800 per year
  • Full-time position
  • 40 hours/week
  • Domestic and international travel required up to 20%
  • Reasonable accommodations for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws
]]>