Lead and grow the system team — set technical direction, mentor engineers, and own hiring — while staying hands-on.
Own the on-prem edge fleet — thousands of in-store servers and cameras across customer networks, reached over resilient VPN/mesh paths (OpenVPN, Tailscale).
Build the internal tools, scripts, and APIs that let the CS team and on-site technicians install and troubleshoot store servers, networks, and cameras on their own — turning cross-team escalations into self-serve fixes.
Own our self-hosted observability platform — a Prometheus-ecosystem stack (Grafana, Mimir, Loki, VMAgent, Vector) across edge and cloud, with SLO dashboards and alerting that feeds automated ticketing and remediation.
Drive the automation that keeps operational load flat as the fleet grows — event-driven auto-ticketing and auto-remediation across Ansible/AWX and serverless AWS (SNS/SQS/Lambda).
Run the core infrastructure and services at the Taipei headquarters — servers and office network, virtualization, Kubernetes, storage, device monitoring, and self-hosted services (container registry, auth, reverse proxy).
Steer the roadmap so systems work enables product velocity — partnering with ML engineers, product managers, customer success, and customer-side IT.
Run the on-call rotation — lead incident response, and turn recurring issues into durable fixes and runbooks.
Own security and compliance operations — the ISO 27001 ISMS, vulnerability management, code security scanning, and disaster-recovery drills.
A track record of leading an infrastructure, platform, or SRE team while staying hands-on in the systems yourself.
5+ years in systems, infrastructure, or DevOps engineering, with strong Linux (Ubuntu) administration and real experience operating fleets at scale.
Strong networking fundamentals — TCP/IP, DNS, VLANs, VPN, and firewalls.
Deep infrastructure-as-code, config-management, and orchestration experience across on-prem and cloud — e.g. Terraform, Pulumi.
Observability fluency — running metrics, logs, and alerting at scale, defining SLOs, and balancing monitoring detail against the storage and cost it drives.
Solid automation and scripting skills — Bash and Python especially, and ideally some Go — you reach for code to remove repetitive operational work.
Excellent communication — clear docs and runbooks, calm incident leadership, and coordination across product, customer success, customers, and third-party IT.
Fluent in Mandarin and English — Mandarin is the team's working language; fluent English is required for daily work with US teams, customers, and IT vendors.
A pragmatic sense of ownership — balancing reliability against delivery velocity, and knowing which problems are worth solving now.
Familiarity with IP cameras/NVRs and the ONVIF protocol, plus real-time video streaming (RTSP, H.264/H.265, MediaMTX).
Storage and NAS operations at scale (ZFS/TrueNAS, Synology, QNAP, RAID/HA).
Hands-on vulnerability management and SAST — e.g. OpenVAS, SonarQube.
Workflow-automation tooling (n8n or similar) and a track record of building low-toil operational systems.
Small team, high ownership, fast feedback from customers — and the operational rigor to make that velocity sustainable. Modern AI tooling — LLMs, coding agents, agent-driven workflows — is a normal part of how we work, and you're encouraged to push on what these tools can do.
============================
Online (Google Meet)
Team Lead Interview (0.5 - 1 hr)
Onsite
Technical Interview (1.5 hrs)
CEO & VP Interview (1.5 hrs)
HR Interview (0.5 hr)
Job details are sourced from the employer's original posting.
Open job postingAbout the company
Berry AI is a company focused on developing AI solutions.