We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Senior Site Reliability Engineer NEX

Patterson-UTI
United States, Texas, Houston
10713 West Sam Houston Parkway North (Show on map)
Aug 14, 2026

Reliability Engineering

* Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform.

* Improve service availability, latency, performance, scalability, and operational resilience.

* Define, implement, and track service-level indicators, service-level objectives, and error budgets.

* Perform capacity planning, performance analysis, and workload forecasting.

* Design and validate disaster recovery, backup, failover, and service-restoration capabilities.

* Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices.

* Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards.

* Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency.

* Identify operational risks and recommend improvements to cloud architecture and service design.

Automation and Platform Engineering

* Build and maintain cloud infrastructure using Terraform or comparable infrastructure-as-code tools.

* Automate repetitive operational activities and systematically identify, measure, and reduce manual toil.

* Build reusable infrastructure modules, deployment patterns, and operational tooling.

* Improve CI/CD pipelines to enable secure, repeatable, and reliable software delivery.

Observability and Incident Management

* Develop actionable alerts that identify meaningful service degradation while reducing alert fatigue and unnecessary operational noise.

* Create and maintain dashboards, runbooks, operational procedures, and troubleshooting documentation.

* Participate in a sustainable on-call rotation supporting production systems.

* Respond to production incidents, coordinate service restoration, and lead incident response when appropriate.

* Facilitate blameless postmortems and identify corrective and preventive actions.

* Use incident and operational data to improve system design, automation, monitoring, and response processes.

Collaboration and Service Ownership

* Partner with software engineering, data engineering, security, and product teams to improve application reliability and production operations.

* Promote shared responsibility for production reliability between application development and platform teams.

* Establish and document reliability standards, operational practices, and reusable engineering patterns.

* Provide technical guidance and coaching on SRE, cloud, Kubernetes, observability, and incident-management practices.

Required Knowledge, Skills, and Abilities

* Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud engineering, production software engineering, or a similar role.

* Experience operating highly available systems in a 24/7 production environment.

* Hands-on experience operating workloads on Google Cloud Platform or another major public cloud platform.

* Strong experience managing compute, networking and data GCP services workloads

* Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.

* Experience building and managing infrastructure with Terraform or a comparable infrastructure-as-code tool.

* Proficiency in Python, Go, Java, or another comparable programming language.

* Experience implementing or operating CI/CD pipelines using GitHub Actions,Azure DevOps, Bitbucket Pipelines, or comparable tools.

* Experience implementing observability using metrics, logs, traces, dashboards, and alerts.

* Experience participating in on-call rotations, responding to incidents, and contributing to postmortems.

* Understanding of SLIs, SLOs, error budgets, and other SRE principles.

* Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud-service layers.

* Ability to communicate effectively with engineering teams, business stakeholders, and operational personnel.

Minimum Qualifications

* Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.

* 3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, or DevOps.

* 3+ years of experience operating production workloads in GCP.

* Ability to understand and communicate in English at a level sufficient to issue, receive, and respond to safety-related and operations-related instructions.

Preferred Qualifications

* Google Cloud and/or Kubernetes certifications.

* Experience supporting data-intensive, streaming, analytics, or event-driven platforms.

* Experience establishing production-readiness, incident-management, or reliability-review processes.

* Experience working in the energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries.

* Experience supporting technology environments that integrate cloud platforms with remote sites, field equipment, industrial systems, or edge computing.

Applied = 0

(web-77cf7d65c7-lg9kd)