Software Engineer- GPU Fabric Observability Baseten Labs IncSoftware Engineer- GPU Fabric ObservabilitySan Francisco, CAAs we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures.
Staff Software Engineer - AI Research Infrastructure Databricks IncStaff Software Engineer - AI Research InfrastructureSan Francisco, CA$190,000–$270,000 / yearYou will design and build services that schedule, orchestrate, and observe large‑scale training and inference experiment workloads across thousands of GPUs, improve our dev tooling and ensure that researchers can iterate quickly without sacrificing reliability, efficiency, or security. As a Staff Software Engineer on the AI Research Infra Team at Databricks, you will: Design and implement infrastructure that supports large‑scale experiments, data processing, and model training (e.g., HPC clusters, GPU fleets, or cloud‑based systems).
Lead, IT Applications Ross Stores IncLead, IT ApplicationsDublin, CA$114,300–$195,000 / yearManage projects throughout the entire IT Project Delivery lifecycle (SDLC), ensuring that requirements are defined accurately, systems are designed / coded or purchased that meet the defined requirements and are executed efficiently. ESSENTIAL FUNCTIONS: Develop partnership, acting as a liaison between technical and business teams to understand, troubleshoot, interpret, and advise on technical questions/issues/projects or business use cases.
Software Engineer - Training Infrastructure BasetenSoftware Engineer - Training InfrastructureSan Francisco, CaliforniaFamiliarity or experience with the open source training stack and frameworks (NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainer) and distributed training techniques (FSDP, DeepSpeed). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
NewPARK AIDE (SEASONAL) (ANGEL ISLAND) State Of CaliforniaPARK AIDE (SEASONAL) (ANGEL ISLAND)CA$17.38–$20.69 / hourPrimary duties include assisting park visitors, staff, and residents with transportation needs by driving a nine-passenger shuttle van around the island; assisting with park visitor education through interpreting the park's natural and cultural resources for visitors; providing assistance to park visitors by providing directions and information about points of interest; cleaning and maintaining vehicles and assisting with cleaning and removing graffiti from structures; deckhand duties on State Park vessels; greeting visitors on public docks and other duties as required. The mission of California State Parks is to provide for the health, inspiration, and education of the people of California by helping to preserve the state's extraordinary biological diversity, protecting its most valued natural and cultural resources, and creating opportunities for high-quality outdoor recreation.
SENIOR UTILITIES ENGINEER (SPECIALIST) State Of CaliforniaSENIOR UTILITIES ENGINEER (SPECIALIST)San Francisco, CA$11,437–$14,315 / yearThe Project Manager will have lead responsibility in conducting complex technical, policy, and/or economic analyses and research to support Administrative Law Judges (ALJ), Commissioners, and Advisors in the permitting of utility infrastructure projects including electric transmission infrastructure (such as transmission lines and substations), and natural gas infrastructure (such as compressor stations and pipelines). Under general direction of the section Program and Project Supervisor, the Senior Utilities Engineer (Specialist) will act as the CEQA Project Manager responsible for completing California Environmental Quality Act (CEQA) environmental review for infrastructure projects, such as electrical and natural gas infrastructure projects.
Senior Site Reliability Engineer Andromeda ClusterSenior Site Reliability EngineerSan Francisco, CaliforniaIncident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast. Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.
Member of Technical Staff - Infrastructure Gimlet LabsMember of Technical Staff - InfrastructureSan Francisco, CaliforniaYou'll work across bare metal, Linux, Kubernetes and cluster schedulers, high-speed networking, observability, and automation to ensure AI workloads can execute efficiently in production. You'll build the operational systems that abstract this complexity, allowing new silicon to become production-ready quickly while ensuring workloads remain reliable, observable, and performant from day one.
Senior Project Manager, Capital Projects Sila NanotechnologiesSenior Project Manager, Capital ProjectsAlameda, CA$151,000–$204,000 / yearYou will serve as a key leader, directly and indirectly supervising specific project team roles—such as Project Engineers, Construction Managers, and Project Coordinators—while collaborating with internal teams and external engineering and construction partners to effectively coordinate, execute, and close out this large-scale capital project. Project Leadership: Direct and lead key phases of large capital projects and construction management including: project planning, site development, permitting, contracts, front-end engineering and detailed design (FEED), documentation, construction, and closeout.
NewMember of Technical Staff (Software Engineer, GPU Cluster Infrastructure) PerplexityMember of Technical Staff (Software Engineer, GPU Cluster Infrastructure)San Francisco, CaliforniaDesign and own the systems that let inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure. Today, our inference engineers and researchers build models while also managing networking, securing capacity, and operating the underlying GPU clusters, responsibilities we want a dedicated platform team to own.
Software Engineer, Back-End SuperhumanSoftware Engineer, Back-EndSan Francisco, CAOur Back-End engineers build the systems that power AI agents across the entire product suite: the scheduling layer that manages getting all your agents to run at the right time in Go, the real-time infrastructure behind Superhuman Mail, the collaboration engine in Superhuman Docs, and the enterprise controls that let organizations deploy AI safely at scale. You'll build the services that make agentic workflows possible - session history, task scheduling, client-server protocols, and the real-time infrastructure that connects our AI layer to users across every surface.
Software Engineer - Training Infrastructure BaseTenSoftware Engineer - Training InfrastructureSan Francisco, CAFamiliarity or experience with the open source training stack and frameworks (NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainer) and distributed training techniques (FSDP, DeepSpeed). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
NewPUBLIC UTILITIES REGULATORY ANALYST V State Of CaliforniaPUBLIC UTILITIES REGULATORY ANALYST VSan Francisco, CA$9,631–$12,054 / yearUnder general direction of the section Program and Project Supervisor, incumbent will have lead responsibility in conducting complex economic, policy and/or technical analyses and research to support Administrative Law Judges (ALJ), Commissioners, and Advisors in the permitting of utility infrastructure projects including electric transmission lines, gas pipelines, telecommunication facilities, and rail projects. The Infrastructure Planning and CEQA Section in the CEQA-FERC Branch plays an important role in reviewing telecommunication projects throughout the State and is responsible for the development of environmental documents in accordance with CEQA and overseeing projects in construction.
Senior UNIX Production Support Engineer (Linux, AutoSys, Shell Scripting, SQL) Pyramid, IncSenior UNIX Production Support Engineer (Linux, AutoSys, Shell Scripting, SQL)Livermore, CA$45–$55 / hourFull timeKey Requirements and Technology Experience: 5+ years of UNIX/Linux administration Production Support experience AutoSys Scheduler Shell Scripting (KSH, Shell) SQL Batch Processing FTP / SFTP / Connect / XCOM UNIX Troubleshooting TCP/IP Networking Excellent analytical and communication skills Oracle Linux Administration OpenShift Splunk Java SAP R/3 AWK SED Mainframe ACF2 Our client is a leading Healthcare Industry and we are currently interviewing to fill this and other similar contract positions. By applying to our jobs you agree to receive calls, AI-generated calls, text messages, or emails from Pyramid Consulting, Inc. and its affiliates, and contracted partners.
Senior Product Manager, Compute Platform Emerald AISenior Product Manager, Compute PlatformOakland, CaliforniaThis is a highly technical product role where you'll work alongside engineering to define platform capabilities, translate complex customer and technical requirements into clear product direction, and drive execution from concept through production. Emerald AI is building the software platform that enables AI infrastructure to intelligently orchestrate compute in response to power availability, grid conditions, and operational constraints.
NewCustomer Reliability Engineer Andromeda ClusterCustomer Reliability EngineerSan Francisco, CaliforniaRemoteWork GPU failures: driver and device-plugin issues, XID errors, thermal throttling, nodes that need cordoning or draining, jobs failing across multiple nodes. Dig into Kubernetes problems like pods stuck pending or crash-looping, node conditions, scheduling failures, resource limits.
Technical Program Manager, Compute AnthropicTechnical Program Manager, ComputeSan Francisco, CAThis research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. You'll join a small, high-impact TPM team and take ownership of critical workstreams across the compute lifecycle, from how supply is procured and brought online, to how capacity is allocated and utilized across teams.
Junior Staging Designer Vesta HomeJunior Staging DesignerSan Francisco, CAAct as the primary point of contact for clients on all staging installations, partnering closely with the sales team and VP of the market to ensure that all projects are a success. Create an invoice for the project, including hours spent on the project and images of the receipts associated with each expenditure, as related to the specific project.
NewKernel Senior Software Engineer, Fuchsia Google LLCKernel Senior Software Engineer, FuchsiaSan Francisco, CAWe"re looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. You will address complex bare-metal challenges such as managing hardware interrupts, avoiding system-wide kernel panics, and optimizing code generation for safe execution, working in close alignment with cross-functional compiler, tools, and security teams.
Senior Site Reliability Engineer - Core Cloud Platform Lambda IncSenior Site Reliability Engineer - Core Cloud PlatformSan Francisco, CAOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda's designated work from home day is currently Tuesday.