Back to all jobs

HPC and Storage Engineering Roles at Stanford Research Computing

HackerNews
Apply NowSign in to track
AI-enhanced for better readability

Stanford Research Computing

Location: Stanford, CA (next to Palo Alto)
Type: Full-time
Work Model: Hybrid / Onsite
Source: Hacker News


About Stanford Research Computing

Stanford Research Computing is a collaboration between University IT and the Vice Provost and Dean of Research. We operate HPC environments for researchers, provide one-time consultations on projects (ranging from software and pipelines to data management and physical building design), and provide contract support for individual Labs, Departments, and Schools.

We are currently hiring for four positions: three hybrid and one onsite.


Open Positions

1. Principal Storage Architect & Team Lead (Hybrid)

This is a technical manager position. You will lead the storage team and set the direction for our large storage environments, including Oak (file storage), Fir (fast scratch for Sherlock), and Elm (object storage on top of tape).

2. Storage Architect or Storage Sysadmin (Hybrid)

You will be responsible for maintaining and expanding Oak, our 20+ Pebibyte Lustre storage environment used by our largest HPC clusters. Depending on experience, you may also assist with Elm (object storage on top of tape).

3. GPU System Engineer (Hybrid)

We are looking for a lead sysadmin for Marlowe—our 1SU NVIDIA DGX H100 SuperPOD with DDN Intelliflash and DDN NFS storage. You will work with the latest AI/ML/Deep Learning/LLM software and frameworks, ensuring they function within an HPC environment.

  • Responsibilities: Keeping the environment up-to-date, working with NVIDIA/DDN vendors, and interacting with users and PIs.
  • More info & Apply

4. HPC Hardware & Infra Sysadmin (ONSITE)

We are looking for a system administrator to help run the hardware and infrastructure for Sherlock, our largest HPC cluster. Sherlock consists of a mix of Intel & AMD x86_64 servers, three Infiniband fabrics, and an Ethernet backbone.

  • Responsibilities: Maintaining, troubleshooting, and improving the hardware and infrastructure of the cluster.
  • More info & Apply

Benefits & Perks

  • Relocation: Relocation incentive provided if you do not currently live in the Bay Area.
  • Transit: Free transit passes provided (depending on location).
  • Time Off: 30+ days off per year (holidays + vacation).
  • Retirement: 403(b) match.
  • Health: Comprehensive healthcare coverage.
  • Additional Info: All benefits are publicly documented at Cardinal at Work.

Note: If you drive, you will be responsible for parking costs for the days you are on-site. There is some on-call requirement around the holidays.


How to Apply

If you have questions, feel free to reply to the original thread or email the poster (contact information available in their profile).

Similar jobs