HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    AN

    Andglobal

    Technology

    Senior SRE

    Ulaanbaatar, MongoliaOn-SiteFull-time5+ yrs experiencePosted 1mo ago
    All Andglobal jobs

    Job description

    As a Senior Site Reliability Engineer (SRE), you will play a key role in designing, building, and maintaining highly available, secure, and scalable infrastructure across AND Global and its subsidiaries. You will collaborate closely with engineering teams to deliver reliable platform solutions, automate operational processes, and ensure system stability while supporting new business initiatives.

    You'll join a collaborative team of four experienced engineers and contribute to shaping infrastructure best practices, operational excellence, and continuous improvement.

    Key Responsibilities:

    • Monitor production systems, application availability, performance, and overall system - health.
    • Build and maintain reliable cloud and Kubernetes infrastructure.
    • Automate infrastructure provisioning, deployment, monitoring, and recovery processes.
    • Investigate production incidents, identify root causes, and implement permanent fixes.
    • Lead incident response and prepare post-incident reports.
    • Work with development teams to improve application reliability and release processes.
    • Define and maintain service-level indicators, service-level objectives, and availability targets.
    • Improve CI/CD pipelines and deployment procedures.
    • Perform capacity planning, performance tuning, and cost optimization.
    • Build dashboards, alerts, monitoring standards, and operational runbooks.
    • Maintain backup, disaster recovery, and high-availability procedures.
    • Identify technical risks, bottlenecks, and areas for continuous improvement.
    • Mentor junior engineers and support engineering best practices.

    Requirements:

    • Bachelor’s degree in computer science, information technology, or equivalent practical experience.
    • At least 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or system administration.
    • Strong Linux administration and troubleshooting skills.
    • Production experience with AWS, Microsoft Azure, or Google Cloud.
    • Strong experience with Kubernetes and Docker.
    • Experience with infrastructure-as-code tools such as Terraform or OpenTofu.
    • Experience with CI/CD tools such as GitLab CI, GitHub Actions, Jenkins, or Argo CD.
    • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, ELK, OpenSearch, Datadog, or OpenTelemetry.
    • Ability to automate tasks using Python, Go, Bash, or another programming language.
    • Good understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, load balancers, VPNs, firewalls, and routing.
    • Experience supporting distributed applications, databases, storage, and messaging systems.
    • Experience with incident management, root-cause analysis, and postmortems.
    • Strong problem-solving, communication, and documentation skills.

    Job details are sourced from the employer's original posting.

    Open job posting
    AN

    About the company

    Andglobal

    Andglobal is a company focused on developing complex distributed systems and designing innovative tech-based solutions for customer needs. They work closely with other development teams, testers, documentation writers, and product management to deliver high-quality products.

    View all Andglobal jobs
    Industry
    Technology
    Open roles
    1

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.