HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    ER

    Era4

    Technology

    Site Reliability Engineer

    United Kingdom: (Occasional office visit required), United KingdomOn-SiteFull-timePosted 1w ago
    All Era4 jobs

    Job description

    Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations

    Role Summary:

    We’re hiring SRE/Platform engineers with an automation bias to help build Era4’s operations capability from the ground up. You’ll turn runbooks, alerts and operational workflows into safe, auditable automation and internal tooling that improves reliability across our AI infrastructure and datacentre platform.

    This is a Platform / SRE role with software engineering, not an AI model-building role. You’ll work closely with operations, platform and engineering teams to reduce manual toil, improve alert quality, and speed up incident response.

    Key Responsibilities:

    • Build Python-based automation for incident triage, runbook execution, and routine operational tasks.
    • Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate workflows.
    • Improve monitoring signal quality through correlation, enrichment, suppression and deduplication.
    • Build internal tools and self-service capabilities such as CLI utilities, ChatOps integrations and dashboards.
    • Maintain version-controlled runbook-as-code and automation libraries.
    • Translate post-incident learnings into better tooling, automation and operational standards.
    • Support safe, auditable automation for higher-risk actions with appropriate approval controls.

    Essential Experience:

    • Experience in SRE, Platform Engineering, or production infrastructure operations.
    • Hands-on experience with observability/monitoring tooling (for example Prometheus, Grafana or similar).
    • Exposure to incident management / on-call and converting manual runbooks into automation.
    • Experience with Python for automation, APIs and integrations.

    Nice To Have:

    • GPU, datacentre or colocation infrastructure experience.
    • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar).
    • ChatOps tooling (Slack or Microsoft Teams bots).
    • OpenTelemetry, logging or distributed tracing experience.
    • DCIM, IPAM or hypervisor-control-plane integrations.
    • Experience with LLM-assisted or agent-based operational automation.

    Why Join Era4:

    You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

    Diversity & Inclusion:

    Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

    Job details are sourced from the employer's original posting.

    Open job posting
    ER

    About the company

    Era4

    Era4 is a company focused on providing innovative solutions in the technology sector.

    View all Era4 jobsera4.com
    Industry
    Technology
    Open roles
    14

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.