---
title: Senior Site Reliability Engineer (SRE) at EPAM Systems
description: We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable 
---

# Senior Site Reliability Engineer (SRE)

**Company:** EPAM Systems  
**Location:** Georgia  
**Posted:** 2026-08-29  
**Apply by:** 2026-10-13

**Tags:** devops, sre

[Apply / View original posting](https://www.linkedin.com/jobs/view/4459459515)

## Job description

We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis. To discover more about Cloud practice at EPAM Georgia, visit this page. Experience the freedom of remote work from anywhere in Georgia, whether from the comfort of your home, our modern offices in Tbilisi and Batumi or a coworking space in Kutaisi. Responsibilities Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling Automate operational toil through scripting and infrastructure-as-code Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort Lead incident response practices including on-call readiness and blameless post-mortems Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements Requirements 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations English Level: B2+ (Upper-Intermediate) or higher Nice to have Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation Experience with Datadog or similar enterprise observability platforms Background in evangelizing best practices and setting standards across engineering teams Exposure to programmatic advertising or adtech platforms We offer We connect like-minded people Delivering innovative solutions to industry leaders, making a global impact Enjoyable working environment, whether it is the vibrant office or the comfort of your own home Opportunity to work abroad for up to two months per year Relocation opportunities within our offices in 55+ countries Corporate and social events We invest in your growth Leadership development, career advising, soft skills and well-being programs Certifications, including GCP, Azure and AWS Unlimited access to EPAM's internal learning database Free English classes with certified teachers We cover it all Participation in the Employee Stock Purchase Plan Monetary bonuses for engaging in the referral program Comprehensive medical & family care package Five trust days per year (sick leave without a medical certificate) Benefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

---

🔔 [Monitor similar jobs on Gurify](https://gurify.com/?utm_source=techgeo&utm_medium=jobboard&utm_campaign=monitor_similar&role=Senior+Site+Reliability+Engineer+%28SRE%29) — get alerted when matching roles are posted in Georgia.

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Senior Site Reliability Engineer (SRE)","description":"<p>We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis. To discover more about Cloud practice at EPAM Georgia, visit this page. Experience the freedom of remote work from anywhere in Georgia, whether from the comfort of your home, our modern offices in Tbilisi and Batumi or a coworking space in Kutaisi. Responsibilities Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling Automate operational toil through scripting and infrastructure-as-code Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort Lead incident response practices including on-call readiness and blameless post-mortems Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements Requirements 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations English Level: B2+ (Upper-Intermediate) or higher Nice to have Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation Experience with Datadog or similar enterprise observability platforms Background in evangelizing best practices and setting standards across engineering teams Exposure to programmatic advertising or adtech platforms We offer We connect like-minded people Delivering innovative solutions to industry leaders, making a global impact Enjoyable working environment, whether it is the vibrant office or the comfort of your own home Opportunity to work abroad for up to two months per year Relocation opportunities within our offices in 55+ countries Corporate and social events We invest in your growth Leadership development, career advising, soft skills and well-being programs Certifications, including GCP, Azure and AWS Unlimited access to EPAM's internal learning database Free English classes with certified teachers We cover it all Participation in the Employee Stock Purchase Plan Monetary bonuses for engaging in the referral program Comprehensive medical & family care package Five trust days per year (sick leave without a medical certificate) Benefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.</p>","identifier":{"@type":"PropertyValue","name":"TechGeo","value":"4459459515"},"url":"https://techgeo.ge/job/senior-site-reliability-engineer-sre-2f7bcd8","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Georgia","addressCountry":"GE"}},"hiringOrganization":{"@type":"Organization","name":"EPAM Systems"},"directApply":false,"datePosted":"2026-08-29","validThrough":"2026-10-13T23:59:59+04:00","jobLocationType":"TELECOMMUTE","applicantLocationRequirements":{"@type":"Country","name":"Georgia"}}
```
