Are you looking to take the next step in your career as an AWS Cloud Migration Engineer? Let's chat and see if we are a good match!
Opportunity:
Virtual Service Operations is searching for an Environment Management Engineer to build, maintain, and troubleshoot a Kubernetes-based environments end to end from cluster and storage administration through CI/CD, GitOps, and cloud infrastructure. You'll be the person teams turn to when something breaks anywhere in the stack, whether that's a pod, a pipeline, a PVC, or a networking issue, and you'll drive the root-cause analysis and documentation to fix it for good.
Skills:
- Experience administering and troubleshooting Kubernetes-based environments, including OpenShift and/or RKE2
- Strong understanding of Kubernetes resources, including Pods, Deployments, StatefulSets, Services, ConfigMaps, Secrets, Namespaces, and Jobs
- Experience managing and troubleshooting persistent storage, including PVs, PVCs, StorageClasses, and volume mounts
- Experience with Ansible, including development and maintenance of playbooks, roles, inventories, and automation workflows
- Experience with Terraform and Infrastructure as Code (IaC)
- Working knowledge of AWS, including EC2, IAM, S3, EBS, VPC/networking, CloudWatch, and associated infrastructure services
- Experience managing and troubleshooting cloud instances, storage, networking, and access controls
- Experience with Argo CD and GitOps-based deployment practices
- Familiarity with container technologies and container registries such as Harbor, Nexus, and Artifactory
- Experience with Apache Pulsar, including management or troubleshooting of topics and messaging resources
- Strong Linux system administration and troubleshooting skill
- Experience with Git and CI/CD tools such as GitLab CI/CD
- Understanding of Kubernetes/OpenShift RBAC, ServiceAccounts, secrets management, and security best practices
- Familiarity with monitoring, logging, and troubleshooting tools such as Prometheus, Grafana,
- OpenSearch/Elasticsearch, or similar platforms
- Strong understanding of networking fundamentals, including DNS, TCP/IP, ingress/egress, load balancing, and service connectivity
- Ability to troubleshoot complex issues across application, container, Kubernetes cluster, cloud infrastructure, storage, and networking layers
- Strong problem-solving skills with the ability to perform root cause analysis and document technical findings and resolutions