Hiring: Site Reliability Engineer
Location: Hyderabad
Work Mode: Work from Office
Work Shift: 24/7
Experience: 3–8 Years
Level: L2/L3 Engineer
About the Role
We are looking for a Site Reliability Engineer – L2/L3 to join our team in Hyderabad. The ideal candidate will be responsible for maintaining the reliability, availability, and performance of production environments while working closely with engineering and operations teams. You serve as an escalation point, drive root cause analysis, and reduce toil through scripting and tooling.
Key Responsibilities
Own L2/L3 incident response, RCA, and post-mortems for production issues.
Monitor system health and maintain SLA/SLO adherence.
Automate operational tasks to eliminate repetitive toil.
Collaborate with dev teams on deployment reliability and capacity planning.
Participate in on-call rotation and maintain runbooks.
Required Skills:
Operating Systems Hands-on with Linux (RHEL/Ubuntu) — system, process management, file systems, performance tuning. Working knowledge of Windows Server and event log analysis.
Cloud Practical experience on AWS / Azure / GCP — compute, storage, IAM, networking, and managed services. Familiarity with Terraform or equivalent IaC tools.
Scripting & Automation Proficiency in Python and Bash for automation, API interaction, and operational tooling. Exposure to Ansible or similar config management is a plus.
Network Troubleshooting (In-Depth) Strong command of TCP/IP internals — handshake lifecycle, connection states (TIME_WAIT, CLOSE_WAIT, SYN_FLOOD), packet flow, and socket behavior. Hands-on with tools like tcpdump, Wireshark, netstat/ss, traceroute, mtr, and dig. Solid understanding of DNS resolution, TLS/SSL negotiation, NAT, firewalls, and routing. Able to diagnose latency, packet loss, port exhaustion, and network-level bottlenecks at the OS and infrastructure layer.
Application & HTTP Troubleshooting Deep understanding of HTTP/HTTPS methods, status codes, headers, and request lifecycle. Comfortable debugging through curl, Postman, access logs, and reverse proxy configs (Nginx / HAProxy).
Observability Experience with Prometheus, Grafana, Datadog, or ELK. Ability to build dashboards, configure meaningful alerts, and trace issues end-to-end.
Good to Have
Kubernetes / Docker experience.
Familiarity with message queues (Kafka, RabbitMQ).
Basic database troubleshooting (MySQL / PostgreSQL / Redis).
ITIL fundamentals and ITSM tools (Jira SM / ServiceNow).
Soft Skills
Strong analytical thinking, clear communication under pressure, and a bias toward automation and continuous improvement.


About Rackspace Technology
We are the multicloud solutions experts. We combine our expertise with the world’s leading technologies — across applications, data and security — to deliver end-to-end solutions. We have a proven record of advising customers based on their business challenges, designing solutions that scale, building and managing those solutions, and optimizing returns into the future. Named a best place to work, year after year according to Fortune, Forbes and Glassdoor, we attract and develop world-class talent. Join us on our mission to embrace technology, empower customers and deliver the future.
More on Rackspace Technology
Though we’re all different, Rackers thrive through our connection to a central goal: to be a valued member of a winning team on an inspiring mission. We bring our whole selves to work every day. And we embrace the notion that unique perspectives fuel innovation and enable us to best serve our customers and communities around the globe. We welcome you to apply today and want you to know that we are committed to offering equal employment opportunity without regard to age, color, disability, gender reassignment or identity or expression, genetic information, marital or civil partner status, pregnancy or maternity status, military or veteran status, nationality, ethnic or national origin, race, religion or belief, sexual orientation, or any legally protected characteristic. If you have a disability or special need that requires accommodation, please let us know.
Job details are sourced from the employer's original posting.
Open job postingAbout the company
Rackspace is a leading provider of IT as a service in today’s multi-cloud world. It delivers expert advice and integrated managed services across applications, data, security and infrastructure, including public and private clouds and managed hosting. Rackspace partners with every leading technology provider, including Alibaba, AWS, Google, Microsoft, OpenStack, Oracle, SAP, and VMware. The company is uniquely positioned to provide unbiased expertise on which technologies will best serve each customer’s needs. Rackspace was named a leader in the 2018 Gartner Magic Quadrant for Public Cloud Infrastructure Managed Service Providers, Worldwide and has been honored by Fortune, Glassdoor and others as one of the best places to work. Based in San Antonio, Texas, Rackspace serves more than 1