This is a two-year contract of employment, inclusive of benefits.
The Academic Research Services team at UCSF is seeking an Storage Systems Engineer (SYS ADM 4) to serve as a technical resource in the design, deployment, and operation of large-scale(multi petabytes) research storage and data infrastructure. This role will work in close partnership with the FAC Storage Lead to support UCSF’s evolving research ecosystem, including Physical/Virtual Compute, CoreHPC, the Research Analysis Environment (RAE), and large institutional storage initiatives.
This position is primarily responsible for architecture, implementation, and lifecycle management for the Facility for Advanced Computing (FAC), storage and systems, including support for large storage environments based on ZFS, NFS v3 and v4 & SAMBA(SMB) NSF-funded infrastructure, and OS Nexus–aligned data platforms. The role ensures seamless integration between storage systems and the Physical/Virtual Compute as well as CoreHPC compute cluster, enabling performant, reliable, and scalable data access for AI, data science, and computational research workloads.
The Storage Systems Engineer will:
Work with the FAC Storage lead to continue supporting the design and evolution of storage architecture on ZFS across on-prem and hybrid environments, including ZFS, NFS, SMB VAST, parallel filesystems, and enterprise storage platforms
Develop and maintain data movement strategies and tooling (e.g., rsync, rclone, Globus, NFS, SMB workflows) to support large-scale data ingestion, migration, and lifecycle management
Ensure tight integration between storage and Physical/Virtual Compute as well as CoreHPC compute cluster systems, optimizing throughput, latency, and reliability for distributed workloads
Support and scale storage systems backing major institutional initiatives (FAC storage(ZFS), OS Nexus integration)
Collaborate closely with DevOps, networking, and security teams to deliver cohesive research infrastructure solutions
Design and implement monitoring, performance tuning, and capacity planning strategies for storage and data systems
Troubleshoot complex issues across storage, networking, and compute boundaries
Participate in system upgrades, migrations, and expansion efforts with minimal disruption to researchers
Provide guidance to researchers on data organization, transfer strategies, and performance optimization
Evaluate and recommend emerging storage technologies and architectures
This role may lead storage-focused projects and contribute to cross-functional initiatives that improve the scalability, usability, and reliability of UCSF’s research computing ecosystem.
Department Overview
Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers’ needs.
The Research Infrastructure team of the Academic Research Service (ARS) focuses on large scale research platform support, high performance computational and storage services for UCSF researchers so they can address complex computational, AI, and data science problems.
%
of time
Essential Function (Yes/No)
Key Responsibilities
(To be completed by Supervisor)
25
Storage Architecture & Infrastructure
Design, deploy, and operate large-scale storage systems, including ZFS, VAST and parallel filesystems.
Define standards for performance, redundancy, and scalability on ZFS filesystems
Lead the evolution of institutional storage platforms, including FAC storage environments.Primarily ZFS
Manage Active Directory integration with storage and compute systems
15
Data Movement & Migration
Architect and execute large-scale data migrations.
Develop and maintain data movement workflows using tools such as rsync, rclone, and Globus.
Optimize data transfer processes across storage and compute environments.
15
Virtual/Physical Compute & HPC Integration & Performance
Integrate storage systems with the CoreHPC compute cluster
Optimize I/O performance for AI, machine learning, and HPC workloads.
Support efficient data access patterns for distributed and scheduled workloads.
Manage & Configure VMWare, Bare Metal Servers
30
Operations & Reliability
Implement monitoring, alerting, and capacity planning for storage systems & operating systems like Linux/Windows
Troubleshoot issues across storage, network, and compute infrastructure.
Perform system maintenance, patching, and lifecycle management.
10
Researcher Enablement
Advise researchers on data workflows and storage best practices.
Support onboarding of projects with large-scale data requirements.
5
Collaboration & Strategy
Collaborate with DevOps, networking, and security teams.
Evaluate and recommend new storage technologies and architectures.
100%
(To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.)REQUIRED QUALIFICATIONS
PREFERRED QUALIFICATIONS
Job details are sourced from the employer's original posting.
Open job postingAbout the company
UCSF is a leading university dedicated to patient care, research, and education.