ML Engineer - Inference & Model Deployment HiringCafeML Engineer - Inference & Model DeploymentCupertino, CaliforniaYou will own the bridge between model development and real user-facing infrastructure: deploying models, optimizing inference latency and throughput, scaling serving systems, and making sure our models run efficiently in production. Implement optimization techniques such as quantization, pruning, batching, caching, efficient attention, and precision trade-offs while preserving model quality.
Machine Learning: Multimodal Foundation Models The Bot CompanyMachine Learning: Multimodal Foundation ModelsSan Francisco, CaliforniaImprove Cross-Modal Reasoning: Research and implement methods to ensure the model doesn't just "associate" modalities but actually reasons through them (e.g., grounding visual physics in kinematic constraints). Ship and Iterate on Real Systems: Integrate models into real robotic stacks, build on robot code to deploy your models, and optimize performance for edge inference.
Senior Software Engineer - Model Performance InferenceSenior Software Engineer - Model PerformanceSan Francisco, CaliforniaYour work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet you.
Applied Research Scientist - Foundation Models Ambient.aiApplied Research Scientist - Foundation ModelsRedwood City, CaliforniaPowered by Ambient Pulsar, the first reasoning Vision-Language Model purpose-built for physical security, our platform seamlessly integrates with existing security cameras and physical access control systems to unify monitoring, access control, threat assessment, response, and investigations through an always-on reasoning layer that augments security operators with superhuman capabilities. The momentum speaks for itself: we doubled new ARR in FY26, we process 200M+ video hours per day, and have delivered results for world-class customers including Cisco, ServiceNow, SentinelOne, TikTok, Bayer, and MoMA.