---
title: Senior Site Reliability Engineer at TradingView
description: TradingView is the world’s largest financial analysis platform with more than 100M users across 180+ countries. We build tools that help traders and investors make informed decisions — from advanced c
---

# Senior Site Reliability Engineer

**Company:** TradingView  
**Location:** Tbilisi, Georgia  
**Posted:** 2026-08-10  
**Apply by:** 2026-09-24

**Tags:** devops, sre

[Apply / View original posting](https://www.linkedin.com/jobs/view/4451440321)

## Job description

TradingView is the world’s largest financial analysis platform with more than 100M users across 180+ countries. We build tools that help traders and investors make informed decisions — from advanced charting and market data to collaboration and publishing features. Our products are used daily by millions of individuals and trusted by companies like Revolut, Binance, and CME Group. We’re continuing to grow and scale our platform, and we’re looking for people who care about product quality, take ownership of their work, and want to build systems used by a global audience. About The Team The team plays a key role in delivering the final product to the production environment and commissioning new services. It works with containerization, automation, orchestration, virtualization, and monitoring systems. Responsibilities Investigate production incidents and drive them through resolution until all consequences are fully addressed. Perform root cause analysis (RCA) and participate in postmortem reviews together with engineering and product teams. Track and drive corrective actions resulting from incidents and postmortem activities. Develop and improve monitoring, alerting, and observability for assigned services. Define and maintain SLI/SLOs, analyze reliability metrics, and monitor service health objectives. Take ownership of SLA compliance for assigned frontend and backend services. Define and implement availability requirements at the service level, in coordination with engineering and product owners. Identify gaps in monitoring and observability and implement improvements to reduce detection and recovery times. Develop and maintain runbooks, troubleshooting guides, and service recovery procedures. Participate in incident validation and escalation reviews. Review and maintain operational and technical documentation. Analyze service performance, resource utilization, and capacity-related risks. Contribute to automation of diagnostics, incident response, and operational workflows. Share operational knowledge and train duty engineers on new tools, procedures, and best practices. Participate in on-call rotations and assist with complex or cross-service incidents. What makes you the perfect fit Experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or in a similar role. Hands-on experience investigating production incidents and performing root cause analysis. Strong understanding of monitoring, alerting, and observability principles. Experience working with metrics, logs, and distributed tracing systems. Knowledge of SLA, SLI, SLO, and error budget concepts. Experience creating and maintaining operational documentation and runbooks. Strong analytical and troubleshooting skills with the ability to drive investigations to actionable outcomes. Understanding of distributed systems and high-load environments. Experience automating operational and repetitive tasks. Will be a plus CKA (Certified Kubernetes Administrator) or/and CKAD (Certified Kubernetes Application Developer) certifications. Experience with Prometheus, Grafana, OpenTelemetry, or similar observability platforms. Experience implementing reliability practices within engineering organizations. Experience with incident management and problem management processes. Understanding of Google Four Golden Signals. Experience using AI-assisted tools for incident investigation and diagnostics. Experience operating large-scale distributed systems. What we offer you Flexible working hours and a hybrid work format Well-equipped offices for focused and collaborative work A global, distributed team of 500+ professionals Learning, mentorship, and long-term career growth Relocation support and private health insurance Performance-based bonuses TradingView Premium access Regular team events and company-wide meetups Join the TradingView team and help us build a product used by millions of traders and investors worldwide. We look forward to hearing from you! TradingView is an equal opportunity employer. We embrace diversity and are dedicated to fostering a diverse and inclusive workplace. Our success is driven by 600+ professionals from 40+ countries who speak nearly 20 languages.

---

🔔 [Monitor similar jobs on Gurify](https://gurify.com/?utm_source=techgeo&utm_medium=jobboard&utm_campaign=monitor_similar&role=Senior+Site+Reliability+Engineer) — get alerted when matching roles are posted in Georgia.

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Senior Site Reliability Engineer","description":"<p>TradingView is the world’s largest financial analysis platform with more than 100M users across 180+ countries. We build tools that help traders and investors make informed decisions — from advanced charting and market data to collaboration and publishing features. Our products are used daily by millions of individuals and trusted by companies like Revolut, Binance, and CME Group. We’re continuing to grow and scale our platform, and we’re looking for people who care about product quality, take ownership of their work, and want to build systems used by a global audience. About The Team The team plays a key role in delivering the final product to the production environment and commissioning new services. It works with containerization, automation, orchestration, virtualization, and monitoring systems. Responsibilities Investigate production incidents and drive them through resolution until all consequences are fully addressed. Perform root cause analysis (RCA) and participate in postmortem reviews together with engineering and product teams. Track and drive corrective actions resulting from incidents and postmortem activities. Develop and improve monitoring, alerting, and observability for assigned services. Define and maintain SLI/SLOs, analyze reliability metrics, and monitor service health objectives. Take ownership of SLA compliance for assigned frontend and backend services. Define and implement availability requirements at the service level, in coordination with engineering and product owners. Identify gaps in monitoring and observability and implement improvements to reduce detection and recovery times. Develop and maintain runbooks, troubleshooting guides, and service recovery procedures. Participate in incident validation and escalation reviews. Review and maintain operational and technical documentation. Analyze service performance, resource utilization, and capacity-related risks. Contribute to automation of diagnostics, incident response, and operational workflows. Share operational knowledge and train duty engineers on new tools, procedures, and best practices. Participate in on-call rotations and assist with complex or cross-service incidents. What makes you the perfect fit Experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or in a similar role. Hands-on experience investigating production incidents and performing root cause analysis. Strong understanding of monitoring, alerting, and observability principles. Experience working with metrics, logs, and distributed tracing systems. Knowledge of SLA, SLI, SLO, and error budget concepts. Experience creating and maintaining operational documentation and runbooks. Strong analytical and troubleshooting skills with the ability to drive investigations to actionable outcomes. Understanding of distributed systems and high-load environments. Experience automating operational and repetitive tasks. Will be a plus CKA (Certified Kubernetes Administrator) or/and CKAD (Certified Kubernetes Application Developer) certifications. Experience with Prometheus, Grafana, OpenTelemetry, or similar observability platforms. Experience implementing reliability practices within engineering organizations. Experience with incident management and problem management processes. Understanding of Google Four Golden Signals. Experience using AI-assisted tools for incident investigation and diagnostics. Experience operating large-scale distributed systems. What we offer you Flexible working hours and a hybrid work format Well-equipped offices for focused and collaborative work A global, distributed team of 500+ professionals Learning, mentorship, and long-term career growth Relocation support and private health insurance Performance-based bonuses TradingView Premium access Regular team events and company-wide meetups Join the TradingView team and help us build a product used by millions of traders and investors worldwide. We look forward to hearing from you! TradingView is an equal opportunity employer. We embrace diversity and are dedicated to fostering a diverse and inclusive workplace. Our success is driven by 600+ professionals from 40+ countries who speak nearly 20 languages.</p>","identifier":{"@type":"PropertyValue","name":"TechGeo","value":"4451440321"},"url":"https://techgeo.ge/job/senior-site-reliability-engineer-3b0557a","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Tbilisi, Georgia","addressCountry":"GE"}},"hiringOrganization":{"@type":"Organization","name":"TradingView"},"directApply":false,"datePosted":"2026-08-10","validThrough":"2026-09-24T23:59:59+04:00"}
```
