HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    GO

    Google

    Technology

    Site Reliability Manager

    Bengaluru, IndiaOn-SiteFull-time5+ yrs experiencePosted 1w ago
    All Google jobs

    Job description

    info_outline
    XIn most instances, this position requires in-person interviews as part of the hiring process.

    Minimum qualifications:

    • Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
    • 5 years of experience building or managing distributed systems or cloud infrastructure, with a focus on Kubernetes.
    • 5 years of experience in people management.
    • Experience with site reliability engineering, system design, distributed computing.

    Preferred qualifications:

    • 5 years of experience in people management, with managing distributed, multi-site teams through engineering managers or tech leads.
    • Experience in Enterprise tooling and technology.
    • Experience in Systems, Applications, and Products (SAP) or other Enterprise Resource Planning (ERP) systems.

    About the job

    Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.

    Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

    Core Enterprise System (CES) SRE is part of Corporate Engineering-Site Reliability Engineering (SRE). We provide SRE support to Enterprise applications within Google, powering key verticals such as Finance, Legal, Supply Chain, and HR. Our mission is to deliver service excellence with engineering, innovation and customer focus and transform Google's enterprise domain.

    Google is an engineering company at heart. We hire people with a broad set of technical skills who are ready to take on some of technology's greatest challenges and make an impact on users around the world. At Google, engineers not only revolutionize search, they routinely work on scalability and storage solutions, large-scale applications and entirely new platforms for developers around the world. From Google Ads to Chrome, Android to YouTube, social to local, Google engineers are changing the world one technological achievement after another.

    Responsibilities

    • Manage a team of 6-10 site reliability engineers supporting Google’s enterprise services.
    • Develop roadmaps, planning, objectives and key results (OKRs) to move forward the maturity of the managed services. Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation and refinement.
    • Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews. Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
    • Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
    • Practice sustainable incident response ensuring services meet their service level objectives.

    Job details are sourced from the employer's original posting.

    Open job posting
    GO

    About the company

    Google

    Google is a multinational technology company focusing on search, artificial intelligence, cloud computing, and online advertising. It develops and provides a wide range of internet-related services and products.

    View all Google jobsabout.google
    Industry
    Technology
    Founded
    1998
    Open roles
    2927

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.