Storage Systems Engineer

UCSF Medical Center

  • San Francisco, CA
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Academic Researchunmatched
    • Analysis Skillsunmatched
    • Artificial Intelligence (AI)unmatched
    • Best Practicesunmatched
    • Biologyunmatched
    • Biomedical Researchunmatched
    • Cancerunmatched
    • Capacity Managementunmatched
    • Capacity Strategyunmatched
    • Capacity and Performance Managementunmatched
    • Communication Skillsunmatched
    • Community and Social Servicesunmatched
    • Computer Engineeringunmatched
    • Computer Networksunmatched
    • Computer Scienceunmatched
    • Computer Securityunmatched
    • Computer Systemsunmatched
    • Cross-Functionalunmatched
    • Customer Support/Serviceunmatched
    • Data Collectionunmatched
    • Data Managementunmatched
    • Data Migrationunmatched
    • Data Scienceunmatched
    • Data Storageunmatched
    • Dell Computersunmatched
    • DevOpsunmatched
    • Diseaseunmatched
    • Diversityunmatched
    • Ecosystemsunmatched
    • Editingunmatched
    • Emerging Technologyunmatched
    • Employment Contractsunmatched
    • Equal Employment Opportunity (EEO)unmatched
    • Facilities Managementunmatched
    • File Systemsunmatched
    • HIPAA (Health Insurance Portability and Accountability Act)unmatched
    • Health Scienceunmatched
    • Healthcareunmatched
    • IBM Product Familyunmatched
    • Identify Issuesunmatched
    • Input/Outputunmatched
    • Large-Scale Systemsunmatched
    • Linux Operating Systemunmatched
    • Machine Learningunmatched
    • Machine Toolunmatched
    • Microsoft Active Directoryunmatched
    • Microsoft Windows Operating Systemunmatched
    • NSF Audio Formatsunmatched
    • NetApp Storage Systemsunmatched
    • Network Administration/Managementunmatched
    • Network Securityunmatched
    • Onboardingunmatched
    • Operating Systemsunmatched
    • Patient Careunmatched
    • People Managementunmatched
    • Performance Analysisunmatched
    • Performance Managementunmatched
    • Performance Tuning/Optimizationunmatched
    • Problem Solving Skillsunmatched
    • Process Improvementunmatched
    • Rsyncunmatched
    • Schedule Developmentunmatched
    • Scientific Researchunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Patchesunmatched
    • Standards Developmentunmatched
    • Standards of Careunmatched
    • Stem Cellsunmatched
    • Storage Architectureunmatched
    • Storage Softwareunmatched
    • Strategic Planningunmatched
    • System Integration (SI)unmatched
    • System Migrationunmatched
    • System Operationsunmatched
    • Systems Administration/Managementunmatched
    • Systems Engineeringunmatched
    • Systems Maintenanceunmatched
    • Team Playerunmatched
    • Technical Writingunmatched
    • Test Automationunmatched
    • Test Plan/Scheduleunmatched
    • Time Managementunmatched
    • U.S. National Institute of Standards and Technology (NIST)unmatched
    • Usability Engineeringunmatched
    • VMWareunmatched
    • Virtual Machine (VM)unmatched
    • Virtualizationunmatched
    • Writing Skillsunmatched

    Description

    This is a two-year contract of employment, inclusive of benefits.

    The Academic Research Services team at UCSF is seeking an Storage Systems Engineer (SYS ADM 4) to serve as a technical resource in the design, deployment, and operation of large-scale research storage and data infrastructure. This role will work in close partnership with the Senior Research DevOps Engineer to support UCSF's evolving research ecosystem, including CoreHPC, the Research Analysis Environment (RAE), and large institutional storage initiatives.

    This position is primarily responsible for architecture, implementation, and lifecycle management for the Facility for Advanced Computing (FAC), storage and systems, including support for large storage environments, NSF-funded infrastructure, and OS Nexus-aligned data platforms. The role ensures seamless integration between storage systems and the CoreHPC compute cluster, enabling performant, reliable, and scalable data access for AI, data science, and computational research workloads.

    The Storage Systems Engineer will:

    • Work with the lead to continue supporting the design and evolution of storage architecture across on-prem and hybrid environments, including VAST, parallel filesystems, and enterprise storage platforms
    • Develop and maintain data movement strategies and tooling (e.g., rsync, rclone, Globus, SMB workflows) to support large-scale data ingestion, migration, and lifecycle management
    • Ensure tight integration between storage and HPC compute systems, optimizing throughput, latency, and reliability for distributed workloads
    • Support and scale storage systems backing major institutional initiatives (FAC storage, OS Nexus integration)
    • Collaborate closely with DevOps, networking, and security teams to deliver cohesive research infrastructure solutions
    • Design and implement monitoring, performance tuning, and capacity planning strategies for storage and data systems
    • Troubleshoot complex issues across storage, networking, and compute boundaries
    • Participate in system upgrades, migrations, and expansion efforts with minimal disruption to researchers
    • Provide guidance to researchers on data organization, transfer strategies, and performance optimization
    • Evaluate and recommend emerging storage technologies and architectures

    This role may lead storage-focused projects and contribute to cross-functional initiatives that improve the scalability, usability, and reliability of UCSF's research computing ecosystem.

    Department Overview

    Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers' needs.

    The Research Infrastructure team of the Academic Research Service (ARS) focuses on large scale research platform support, high performance computational and storage services for UCSF researchers so they can address complex computational, AI, and data science problems.

    About UCSF

    The University of California, San Francisco (UCSF) is a leading university dedicated to promoting health worldwide through advanced biomedical research, graduate-level education in the life sciences and health professions, and excellence in patient care. It is the only campus in the 10-campus UC system dedicated exclusively to the health sciences. We bring together the world's leading experts in nearly every area of health. We are home to five Nobel laureates who have advanced the understanding of cancer, neurodegenerative diseases, aging and stem cells.

    Pride Values

    UCSF is a diverse community made of people with many skills and talents. We seek candidates whose work experience or community service has prepared them to contribute to our commitment to professionalism, respect, integrity, diversity and excellence - also known as our PRIDE values.

    In addition to our PRIDE values, UCSF is committed to equity - both in how we deliver care as well as our workforce. We are committed to building a broadly diverse community, nurturing a culture that is welcoming and supportive, and engaging diverse ideas for the provision of culturally competent education, discovery, and patient care. Additional information about UCSF is available here.

    Join us to find a rewarding career contributing to improving healthcare worldwide.

    Equal Employment Opportunity

    The University of California is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, protected veteran status, or other protected status under state or federal law.

    Salary Information

    The final salary and offer components are subject to additional approvals based on UC policy.

    Your placement within the salary range is dependent on a number of factors including your work experience and internal equity within this position classification at UCSF. For positions that are represented by a labor union, placement within the salary range will be guided by the rules in the collective bargaining agreement.

    To learn more about the benefits of working at UCSF, including total compensation, please visit: https://ucnet.universityofcalifornia.edu/compensation-and-benefits/index.html

    REQUIRED QUALIFICATIONS

    • Bachelor''s degree in a related area, such as computer science or engineering, and 6+ years of experience with storage infrastructure support and management, or 10+ years of related experience with large-scale storage systems
    • Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) InfiniBand-based clusters
    • Strong knowledge of ZFS, high-performance parallel filesystems, and storage such as GPFS, Lustre, Vast, DDN, etc
    • Advanced knowledge of computer security best practices and policies, including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPAA, or IS-3 requirements
    • Knowledge of HPC job scheduler system design and operation, such as SLURM or PBS,
    • Ability to elicit and communicate technical and non-technical information in a clear and concise manner.
    • Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines.
    • Understanding of system performance monitoring and actions that can be taken to improve or correct performance.
    • Demonstrated advanced knowledge, skills, and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems.
    • Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration.
    • Demonstrated testing and test planning skills. Demonstrated ability to create automated testing.
    • Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner

    PREFERRED QUALIFICATIONS

    • Expert knowledge of Virtual Machines, Bare Metal Servers & HPC systems infrastructure design
    • Knowledge of the design, development and application of technology and systems to meet business needs.
    • General knowledge of other areas of IT. E.g., Active Directory, Domain Controllers, Network Infrastructure.
    • Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions.
    • Professional certification in enterprise storage technologies (e.g., NetApp, Dell EMC PowerScale, IBM Storage Scale, VAST, Pure Storage)

    %

    of time

    Essential Function (Yes/No)

    Key Responsibilities

    (To be completed by Supervisor)

    25

    Storage Architecture & Infrastructure

    Design, deploy, and operate large-scale storage systems, including ZFS, VAST and parallel filesystems.

    Define standards for performance, redundancy, and scalability on ZFS filesystems

    Lead the evolution of institutional storage platforms, including FAC storage environments.Primarily ZFS

    Manage Active Directory integration with storage and compute systems

    15

    Data Movement & Migration

    Architect and execute large-scale data migrations.

    Develop and maintain data movement workflows using tools such as rsync, rclone, and Globus.

    Optimize data transfer processes across storage and compute environments.

    15

    Virtual/Physical Compute & HPC Integration & Performance

    Integrate storage systems with the CoreHPC compute cluster

    Optimize I/O performance for AI, machine learning, and HPC workloads.

    Support efficient data access patterns for distributed and scheduled workloads.

    Manage & Configure VMWare, Bare Metal Servers

    30

    Operations & Reliability

    Implement monitoring, alerting, and capacity planning for storage systems & operating systems like Linux/Windows

    Troubleshoot issues across storage, network, and compute infrastructure.

    Perform system maintenance, patching, and lifecycle management.

    10

    Researcher Enablement

    Advise researchers on data workflows and storage best practices.

    Support onboarding of projects with large-scale data requirements.

    5

    Collaboration & Strategy

    Collaborate with DevOps, networking, and security teams.

    Evaluate and recommend new storage technologies and architectures.

    100%

    (To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.)

    %

    of time

    Essential Function (Yes/No)

    Key Responsibilities

    (To be completed by Supervisor)

    25

    Storage Architecture & Infrastructure

    Design, deploy, and operate large-scale storage systems, including ZFS, VAST and parallel filesystems.

    Define standards for performance, redundancy, and scalability on ZFS filesystems

    Lead the evolution of institutional storage platforms, including FAC storage environments.Primarily ZFS

    Manage Active Directory integration with storage and compute systems

    15

    Data Movement & Migration

    Architect and execute large-scale data migrations.

    Develop and maintain data movement workflows using tools such as rsync, rclone, and Globus.

    Optimize data transfer processes across storage and compute environments.

    15

    Virtual/Physical Compute & HPC Integration & Performance

    Integrate storage systems with the CoreHPC compute cluster

    Optimize I/O performance for AI, machine learning, and HPC workloads.

    Support efficient data access patterns for distributed and scheduled workloads.

    Manage & Configure VMWare, Bare Metal Servers

    30

    Operations & Reliability

    Implement monitoring, alerting, and capacity planning for storage systems & operating systems like Linux/Windows

    Troubleshoot issues across storage, network, and compute infrastructure.

    Perform system maintenance, patching, and lifecycle management.

    10

    Researcher Enablement

    Advise researchers on data workflows and storage best practices.

    Support onboarding of projects with large-scale data requirements.

    5

    Collaboration & Strategy

    Collaborate with DevOps, networking, and security teams.

    Evaluate and recommend new storage technologies and architectures.

    100%

    (To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.)

    Numbers & Facts

    LocationSan Francisco, CA

    Similar Jobs

    See more jobs