Embedding MLOps

interpretai.tech
  • San Francisco, California
    30+ days ago

    Job Description

    Embedding Cluster Engineer: Cloud Infrastructure Specialist

    Summary

    We are seeking an experienced Cluster Infrastructure Engineer to design, implement, and maintain our vector embedding infrastructure on cloud platforms and support our distributed training platform on the cloud. In this role, you will be responsible for creating a scalable, high-performance system that supports both the training of embedding models and efficient inference workflows. The ideal candidate will combine expertise in machine learning infrastructure with strong system design skills to build robust embedding systems that power our AI applications, including vector searches, recommendations, and SOTA model inference for foundation models.

    Key Responsibilities

    • Design and architect a multi-tenanted, cloud-native embedding infrastructure that supports both training and inference workloads.
    • Build scalable vector search and embedding generation services that handle high throughput, maintain low latency, and measure index time.
    • Implement fault-tolerant, high-performance systems for serving embedding models at scale.
    • Develop infrastructure automation using containerization, orchestration, and infrastructure-as-code practices.
    • Optimize embedding storage, indexing, and retrieval systems for performance and cost efficiency.
    • Design and implement robust monitoring and observability solutions to ensure system health and performance.
    • Collaborate with ML engineers and data scientists to understand and support embedding-related workloads.
    • Create scalable pipelines for generating and updating embeddings from various data sources (text, images, audio).
    • Implement security best practices and ensure compliance with data protection requirements3
    • Lead technical discussions regarding vector database architecture and performance optimization.
    • Support hiring efforts for building the core infrastructure.

    Requirements

    • Bachelor's degree in Computer Science, Engineering, or related technical field (Master's preferred).
    • 2+ years of experience building large-scale, high-performance backend systems9
    • 2+ years of experience with cloud platforms (AWS, GCP, or Azure) and infrastructure-as-code tools39
    • Strong proficiency in at least one programming language such as Python, Go, C++, or Java.
    • Experience with containerization and orchestration technologies (Docker, Kubernetes).
    • Knowledge of distributed systems concepts like sharding, replication, and consensus algorithms.
    • Demonstrated experience with database systems, search technologies, or AI/ML systems8
    • Understanding of embedding techniques and their applications for text, images, or other data types.
    • Experience with memory management, networking, and troubleshooting distributed systems.
    • Proven ability to solve complex problems independently in fast-moving environments.

    Preferred Qualifications

    • Experience with vector databases and similarity search systems (FAISS, Pinecone, Milvus, etc.).
    • Knowledge of ML frameworks (PyTorch, TensorFlow) and model optimization techniques.
    • Experience with SOTA embedding models from a variety of domains (CV, LLMs, etc.).
    • Understanding of embedding evaluation metrics and quality assessment techniques.
    • Experience with high-performance computing (HPC) environments.
    • Background in designing systems for machine learning training and/or inference workloads.
    • Knowledge of data sharding strategies, parallel processing, and memory optimization techniques.
    • Experience with real-time streaming data processing frameworks.
    • Contributions to open-source projects related to embeddings or machine learning infrastructure.

    Impact of This Role

    As an Embedding Cluster Engineer, you will build the foundation for our AI capabilities. Your work will enable our company to efficiently process unstructured data, power advanced search capabilities, enhance recommendation systems, and improve the accuracy of our introspection platform for our customers. Your infrastructure will be critical to both the training and deployment of our embedding models, directly impacting the speed and quality of our AI services.

    About Our Team

    Our infrastructure team is focused on building scalable, reliable systems that powers data instrospection and data curation platforms. We value collaboration, innovation, and a strong focus on production excellence. You'll work closely with machine learning engineers, data scientists, and other infrastructure specialists to create systems that turn cutting-edge research into production-ready services.

    Technical Environment

    You'll be working with technologies including:
    • Cloud environments (AWS/GCP/Azure)
    • Kubernetes for container orchestration
    • Vector databases and search engines
    • Distributed computing frameworks
    • ML model serving platforms
    • Infrastructure automation tools
    • Monitoring and observability solutions

    Numbers & Facts

    LocationSan Francisco, California
    Websiteinterpretai.tech

    Skills

    • Algorithmsunmatched
    • Amazon Web Services (AWS)unmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Best Practicesunmatched
    • C++ Programming Languageunmatched
    • Cloud Computingunmatched
    • Computer Scienceunmatched
    • Data Partitioningunmatched
    • Data Processingunmatched
    • Data Scienceunmatched
    • Database Architectureunmatched
    • Database Technologyunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • GCP (Good Clinical Practices)unmatched
    • Go Programming Language (Golang)unmatched
    • High Throughputunmatched
    • Image Editorsunmatched
    • Information/Data Security (InfoSec)unmatched
    • Javaunmatched
    • Machine Learningunmatched
    • Maintain Complianceunmatched
    • Memory Hardwareunmatched
    • Memory Managementunmatched
    • Microsoft Windows Azureunmatched
    • Network Administration/Managementunmatched
    • Open Sourceunmatched
    • Parallel Computingunmatched
    • Performance Tuning/Optimizationunmatched
    • Problem Solving Skillsunmatched
    • Programming Languagesunmatched
    • Python Programming/Scripting Languageunmatched
    • QoS (Quality of Service)unmatched
    • Quality Metricsunmatched
    • Replication and Remote Mirroringunmatched
    • Scalable System Developmentunmatched
    • Search Enginesunmatched
    • Search Technologyunmatched
    • Software Engineeringunmatched
    • Technical Leadershipunmatched
    • Unstructured Dataunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder