Senior/Staff Virtualization Engineer

Fal.ai Inc

San Francisco, CA

Apply

JOB DETAILS

SALARY

$180,000–$250,000 Per Year

SKILLS

ARP (Address Resolution Protocol), Ansible, Artificial Intelligence (AI), Automation, BGP, Bash Scripting, CPU (Central Processing Unit), Cloud Computing, Communication Skills, Computer Systems, Debugging Skills, Dental Insurance, GPU (Graphics Processing Unit), IPsec (IP Security), K Virtual Machine (KVM), Linux Operating System, Network Configuration Management, Network Design, Network Monitoring, Network Routing, Operating Systems, Python Programming/Scripting Language, VLAN (Virtual Local Area Network), Virtual Machine (VM), Virtualization, Vision Plan, Wireshark (Ethereal), tcpdump

LOCATION

San Francisco, CA

POSTED

30+ days ago

You build the custom compute environments we deliver to customers - bare metal or virtual machines with GPU passthrough, dedicated Kubernetes clusters, and the networking that ties them together. You work across the full stack from Linux image building to overlay network design to cluster bootstrapping.

Key responsibilities

Build and deliver custom environments with excellent GPU performance for customer workloads
Leverage AI to an extreme level to automate provisioning, alerting and recovery
Provision and configure dedicated Kubernetes clusters tailored to customer requirements
Design and implement overlay networking (VLAN, VXLAN) and routing configurations (ECMP, BGP) and tunnels (strongSwan, IPSEC) for tenant isolation and performance
Build and maintain Linux images
Set up network monitoring and diagnostics for customer environments
Automate the end-to-end lifecycle of customer compute environments: creation, configuration, validation, and teardown

Requirements

5+ years experience with Linux virtualization: KVM/QEMU, libvirt, VFIO device passthrough, hugepages, NUMA, CPU pinning
Strong networking fundamentals: VXLAN, VLAN, ECMP, BGP, ARP, and the ability to debug packet-level issues (tcpdump, Wireshark)
Production experience building and operating Kubernetes clusters on bare metal (MetalLB)
Proficiency with Linux image building and OS provisioning (kickstart, cloud-init, PXE/iPXE)
Proficiency in Python, Bash, Ansible and Terraform
Deep experience with NVIDIA GPUs: drivers, MIG, container runtimes (nvidia-container-toolkit), InfiniBand, RDMA/RoCEv2 and GPUDirect for high-performance AI networking
Excellent communication and ability to drive technical decisions across teams
Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to have

Experience with SR-IOV, DPDK, or other high-performance networking technologies
Experience with shared network storage (Ceph, Lustre, Weka)
Experience with network automation tools (Netbox, Nautobot, Nornir)

Compensation

$180,000-250,000 plus equity + benefits

Location

San Francisco, CA

What we offer at fal

Interesting and challenging work
A lot of learning and growth opportunities
We are currently hiring in downtown San Francisco.
We offer visa sponsorship and will help you relocate to San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsites

About the Company

Fal.ai Inc

Resume Resources

Free Resume Templates Free Resume Builder