About us
Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.
We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.
We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.
About the role
As a Member of Technical Staff, you will build the inference systems that execute models end-to-end in production.
Your work will shape how workloads are scheduled, how memory is managed, and how the platform balances latency, throughput, and resource utilization.
You will work across model serving, batching, scheduling, concurrency, KV cache management, and memory placement. You will help bring up models on novel hardware. You will support new model architectures and inference techniques, improve performance under real production workloads, and partner with compiler, kernel, networking, and distributed systems engineers to optimize the full execution path.
What success looks like
In the first 12-18 months, you will:
Improve the latency, throughput, and efficiency of production inference workloads
Design execution strategies across batching, scheduling, concurrency, and resource utilization
Improve KV cache management, memory efficiency, and execution under load
Enable new models, accelerator architectures, and inference techniques to run efficiently in production
You may be a good fit if you have
Strong software engineering fundamentals
Experience building or operating ML inference or model serving systems
Comfort reasoning about performance, memory usage, and system behavior under load
Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.
Strong candidates may also have
Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems
Deep understanding of modern model architectures and attention mechanisms
Experience with batching, scheduling, and concurrency control in inference systems
Familiarity with KV cache management and memory placement strategies
Experience profiling and tuning latency- and throughput-critical systems
Software development experience in Python and C++
Why join now?
Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.
Solve hard problems.
Own meaningful work.
Build for production.
Help define what’s next.
Agency Policy:Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.
| Location | San Francisco, California |
| Website | https://www.gimletlabs.ai/ |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder