Software Engineer - Infrastructure

Emergent Labs
  • San Francisco, California
  • Autofill and Review
30+ days ago

Job Description

Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications.

Since our public launch, we've crossed $100M in ARR and grown to over 10M users across 190+ countries, who have built 12M+ applications on Emergent. We're backed by Creaegis, Claypond, Sentinel Global, Khosla Ventures, SoftBank, Google, Lightspeed, Prosus, Together, and Y Combinator.

We're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox.

We're hiring builders who want ownership, speed, and impact at global scale.

What You'll Do:

Platform and Infrastructure

  • Maintain stability of our platform consisting of distributed microservices closely interacting with Kubernetes and cloud providers (GCP, AWS)
  • Manage Kubernetes workloads with ArgoCD (GitOps), deploy, monitor, and troubleshoot application syncs, resource trees, and rollouts
  • Debug and resolve complex Kubernetes issues across clusters
  • Manage CDN and edge infrastructure (Cloudflare) for performance, caching, and traffic management
  • Automate infrastructure lifecycle operations and workflows

Observability and Incident Response

  • Own the observability stack: Grafana (dashboards, Loki logs, Prometheus metrics) and New Relic (APM, golden metrics, transaction analysis)
  • Enhance monitoring, alerting, and distributed tracing across services
  • Participate in on-call rotation via PagerDuty, handle incident response, and perform root cause analysis
  • Proactively identify reliability risks before they become incidents

AI Agent Infrastructure

  • Support the platform that runs AI agent workloads including job scheduling, trajectory tracking, environment provisioning, deployments, and cost attribution
  • Develop Kubernetes controllers and operators to extend platform capabilities for agent orchestration

Collaboration and Internal Tooling

  • Work closely with product and backend teams to ensure platform scalability and reliability
  • Build internal tools, automate workflows, and integrate systems to improve team productivity
  • Stay current with Kubernetes releases, CNCF ecosystem updates, and cloud-native best practices

What We're Looking For:

Core Requirements

  • 3+ years of software/platform engineering experience with production systems
  • Strong proficiency in Go or Python, you write production code in at least one daily
  • Hands-on experience building and deploying services on Kubernetes, not just YAML, you've developed something that runs on K8s
  • Experience with GitOps tooling (ArgoCD, Flux, or similar)

Systems Fundamentals

  • Strong networking and DNS fundamentals: TCP/IP, HTTP, load balancing, DNS resolution, TLS, and debugging connectivity issues
  • Solid Linux/OS fundamentals: process management, filesystem, memory, systemd, and comfortable debugging with tools like strace, tcpdump, and netstat

Data and Messaging Infrastructure

  • Relational databases: experience with PostgreSQL, MySQL, or similar; indexing, query optimization, replication, and backup/restore procedures
  • NoSQL databases: familiarity with MongoDB, DynamoDB, Redis, or similar for document/key-value workloads
  • Caching: experience with Redis, Memcached, or similar for application and infrastructure-level caching
  • Message queues and streaming: hands-on with Kafka, SQS, RabbitMQ, or similar for event-driven architectures
  • Strong SQL skills for debugging and operational queries

Infrastructure and Observability

  • Comfortable with the CNCF ecosystem: Helm, Kustomize, cert-manager, Ingress controllers, CNI/CSI interfaces
  • Hands-on with at least one observability stack (Grafana/Prometheus/Loki, New Relic, Datadog, or similar)
  • Familiarity with GCP and/or AWS: managed Kubernetes (GKE/EKS), networking, IAM, storage, and cloud-native services (SES, SQS, S3, etc.)
  • Experience with CDN/edge platforms (Cloudflare, CloudFront, or similar)

Good to Have:

  • Experience building Kubernetes Operators (kubebuilder, operator-sdk, or controller-runtime)
  • Experience tuning Kubernetes core components (API server, kubelet, scheduler)
  • Familiarity with AI/LLM infrastructure: token management, cost tracking, agent orchestration
  • Experience with CI/CD pipelines (GitHub Actions, automated testing, deployment pipelines)
  • Infrastructure as Code experience (Terraform, Pulumi, or similar)
  • Previous work on large-scale distributed systems or platform-as-a-service
  • Startup experience, you thrive in fast-paced, ambiguous environments

Who You Are:

  • A generalist who can context-switch between debugging a K8s deployment, setting up a Grafana alert, and configuring CDN rules, all in the same day
  • You enjoy solving complex infrastructure challenges and automating away toil
  • You dig deep, when something breaks, you find the root cause, not just the workaround
  • You communicate clearly and can collaborate effectively in a fast-moving, distributed team

Tech Stack:We don't require previous experience with our entire stack, but enthusiasm for learning is key: Go, Python, Kubernetes, ArgoCD, Helm, GCP, AWS, Cloudflare, Grafana, Prometheus, Loki, New Relic, PagerDuty, PostgreSQL, MongoDB, Redis, Kafka, and GitHub.Benefits and Perks:

  • 401(k)
  • Health, dental, and vision insurance
  • Unlimited Paid Time Off: take the time you need to recharge and come back refreshed
  • Flexible Working Hours: work arrangements that fit your life and commitments

Let's build the future of software together.

Numbers & Facts

LocationSan Francisco, California

Skills

  • Amazon CloudFrontunmatched
  • Amazon Simple Storage Service (S3)unmatched
  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Apache Kafkaunmatched
  • Application Programming Interface (API)unmatched
  • Artificial Intelligence (AI)unmatched
  • Best Practicesunmatched
  • Cachingunmatched
  • Cloud Computingunmatched
  • Cloud Storageunmatched
  • Communication Skillsunmatched
  • Content Delivery Network (CDN)unmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • DNS (Domain Name System)unmatched
  • DNS Name Resolutionunmatched
  • Debugging Skillsunmatched
  • Debugging Toolsunmatched
  • Dental Insuranceunmatched
  • Distributed Computingunmatched
  • Ecosystemsunmatched
  • Engineeringunmatched
  • File Systemsunmatched
  • Financial Controlunmatched
  • GCP (Good Clinical Practices)unmatched
  • GitHubunmatched
  • HTTP (HyperText Transport Protocol)unmatched
  • Identify Issuesunmatched
  • Incident Responseunmatched
  • Large-Scale Systemsunmatched
  • Linux Operating Systemunmatched
  • Load Balancingunmatched
  • Machine Toolunmatched
  • Memory Hardwareunmatched
  • Messaging Middlewareunmatched
  • Metricsunmatched
  • Microservicesunmatched
  • MongoDBunmatched
  • MySQLunmatched
  • NoSQLunmatched
  • On Callunmatched
  • PostgreSQLunmatched
  • Problem Solving Skillsunmatched
  • Process Managementunmatched
  • Production Systemsunmatched
  • Productivity Managementunmatched
  • Python Programming/Scripting Languageunmatched
  • RabbitMQunmatched
  • Redisunmatched
  • Relational Databases (RDBMS)unmatched
  • Reporting Dashboardsunmatched
  • Risk Analysisunmatched
  • Root Cause Analysisunmatched
  • SQL (Structured Query Language)unmatched
  • SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
  • Schedule Developmentunmatched
  • Simple Queue Service (SQS)unmatched
  • Software Engineeringunmatched
  • Startupunmatched
  • System Integration (SI)unmatched
  • Systems Engineeringunmatched
  • TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
  • Test Automationunmatched
  • Vision Planunmatched
  • tcpdumpunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder