We’re looking for a Senior Software Engineer to take ownership of the production systems and internal data platform that power our Machine Learning (ML) and Operations Research (OR) work.
This is a hands-on, high-ownership role at the intersection of backend engineering, cloud infrastructure, and data engineering. You’ll work closely with our ML and OR specialists to ensure models and optimization solutions can move reliably from experimentation into production.
You’ll inherit existing production systems and have the opportunity to improve and evolve them over time—from architecture and infrastructure to deployment, observability, and developer tooling. As our needs grow, you’ll also help shape the roadmap for our internal data and ML platform.
Because we’re a small team, you’ll have meaningful autonomy. You’ll be the primary owner of these systems, collaborate directly with technical specialists and product teams, and have significant influence over architecture, tooling, and engineering practices.
Own, operate, and improve backend services running in Azure, including serverless services, batch workloads, and ML inference endpoints.
Manage deployments and reliability across environments, including CI/CD, monitoring, alerting, incident response, and operational runbooks.
Operate and optimize cloud compute environments, including autoscaling, container images, identity, and resource management.
Improve the reliability, scalability, and maintainability of existing production systems over time.
Design and build reliable ETL/ELT pipelines that transform data from relational and document databases into analysis-ready datasets.
Help develop our lakehouse-style analytical layer and the infrastructure that supports it.
Build and maintain infrastructure-as-code across environments using tools such as Terraform or Bicep.
Manage cloud infrastructure including storage, application hosting, identity and access management, Key Vault, and cost optimization.
Implement monitoring, logging, data-quality checks, and freshness alerting across data workflows.
Ensure data is handled securely through appropriate access controls, secrets management, and responsible treatment of sensitive information.
Build new internal services, APIs, and developer tooling as the team's needs evolve.
Build the infrastructure and tooling our ML and OR specialists need for experimentation, deployment, evaluation, and reproducibility.
Turn research prototypes into reliable, production-ready services and workflows.
Own model packaging, versioning, deployment, and CI processes for ML and OR codebases.
Build automated evaluation and benchmarking pipelines to monitor model performance, drift, and system reliability.
Partner with Data Scientists and OR specialists to run and operationalize experiments.
Establish strong practices around code quality, automated testing, version control, and CI/CD.
Conduct peer code reviews and help teammates adopt scalable engineering practices.
Improve existing systems incrementally rather than rebuilding for the sake of rebuilding.
Help ensure our codebases remain maintainable and releasable as the team and platform grow.
Work closely with Data Science, Operations Research, Product, and Engineering to integrate ML and optimization solutions into our products.
Contribute to technical design discussions and decisions around architecture, scalability, reliability, and performance.
Translate technical and business needs into pragmatic engineering solutions.
Bachelor’s or Master’s degree in Computer Science, Software Engineering, Data Engineering, or a related field, or equivalent practical experience.
Strong programming skills in Python and SQL.
Strong understanding of APIs, backend service design, and distributed systems.
Experience building and operating production data pipelines end to end, including concepts such as retries, idempotency, backfills, orchestration, and freshness monitoring.
Hands-on experience designing and operating production services in Azure or another major cloud platform, with an interest in going deep on Azure.
Experience with cloud infrastructure concepts such as serverless and batch compute, object storage, identity and access management, and monitoring.
Experience with infrastructure-as-code tools such as Terraform, Bicep, or ARM.
Strong knowledge of Git, CI/CD, automated testing, and modern software engineering practices.
Comfort working with ML or OR codebases and model artifacts—you don't need to be the person building the models, but you should be comfortable reading, running, packaging, and deploying them.
Experience owning live production systems and confidence taking over an existing codebase, understanding it, and improving it over time.
You don’t need to check every box, but experience with any of the following would be valuable:
Delta Lake, Parquet, or lakehouse architectures.
DuckDB, Polars, or dbt.
Azure Machine Learning, MLflow, or DVC.
Durable Functions or workflow orchestration tools such as Airflow, Dagster, or Prefect.
Kubernetes or other technologies supporting large-scale distributed workloads.
Azure data governance and security practices.
DevOps or SRE practices related to observability, reliability, and performance.
Optimization, logistics, transportation, or large-scale ML systems.
Job details are sourced from the employer's original posting.
Open job postingAbout the company
GoMaterials is a platform that connects landscape contractors with wholesale nurseries and growers. They aim to streamline the sourcing and procurement process for plant materials.