Role Description:
We are seeking a highly motivated and skilled Site Reliability Engineer (SRE) to join our Crew IT team in DTH Bangalore. The SRE will be responsible for ensuring the reliability, availability, scalability, security, and performance of production systems and services. The ideal candidate will combine software engineering expertise with infrastructure and operations knowledge to build automation, improve system resilience, and support mission-critical applications.
Responsibilities:
Reliability & Operations
· Ensure high availability, reliability, and performance of production services.
· Monitor application and infrastructure health using observability and monitoring tools.
· Participate in incident management, root cause analysis (RCA), and post-incident reviews.
· Drive continuous improvements to reduce operational toil and prevent recurring issues.
· Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
Automation & Engineering
· Develop and maintain automation solutions to streamline operational processes.
· Build and enhance CI/CD pipelines to support efficient and reliable software delivery.
· Implement Infrastructure as Code (IaC) practices using tools such as Terraform, ARM/Bicep, or similar technologies.
· Automate deployment, scaling, patching, monitoring, and recovery procedures.
Cloud & Infrastructure Management
· Manage and optimize cloud infrastructure in AWS.
· Support containerized applications and orchestration platforms such as Kubernetes (AKS).
· Ensure effective capacity planning and resource utilization.
· Maintain backup, disaster recovery, and business continuity processes.
Monitoring & Observability
· Implement and maintain monitoring, logging, alerting, and tracing solutions.
· Analyze performance trends and proactively identify system bottlenecks.
· Improve observability across applications and infrastructure.
Security & Compliance
· Collaborate with security teams to implement security best practices.
· Support vulnerability remediation and infrastructure hardening activities.
· Ensure compliance with organizational and regulatory requirements.
Collaboration
· Partner with software engineering teams to improve application reliability and operability.
· Participate in architecture reviews and production readiness assessments.
· Mentor team members on operational excellence and reliability best practices.
Job details are sourced from the employer's original posting.
Open job postingAbout the company
Delta Global Technology Hub is a company focused on technology solutions and innovation.