Software Engineer, Infrastructure

Flexcompute Inc
  • Watertown, MA
    8 days ago

    Job Description

    Flexcompute is a cutting-edge technology startup that specializes in ultra-fast simulation technology. Our products are utilized by companies in designing and optimizing technology products, with applications ranging from designing airplanes and cars to wind turbines and quantum computing chips. Our customer base includes both household names and startups in emerging industries. Our company was founded by world-renowned leaders in simulation technology from Stanford University and MIT. Backed by top VC firms, we are poised to disrupt the billion-dollar engineering simulation industry with our fast-growing trajectory.

    Tidy3D is a GPU-accelerated electromagnetic simulation product delivered as a cloud service. Behind the solver sits the platform that makes everything work: the Python client, the API, the job submission and scheduling layer, and the web application.

    We are hiring a Software Engineer for the Tidy3D infrastructure team to own that platform. You will build new capability, keep the existing system healthy, and run the release and deployment process.

    Responsibilities

    Design and build backend services for the Tidy3D platform.

    • Develop and operate the control plane: the APIs, services, and data model behind task submission, job state, and result delivery.
    • Build scheduling and resource management for simulation jobs across heterogeneous GPU capacity.
    • Handle the operational concerns that come with a multi-tenant product: authentication, authorization, usage metering, and quota enforcement.

    Help with the deployment of our products to customers.

    • Manage packaging and release to PyPI, and keep client and backend versions compatible across a long tail of installed versions.
    • Own the release pipeline end to end: versioning, CI/CD, staged rollout, and rollback.
    • Standardize deployment patterns so the same product ships to our cloud, to customer-managed cloud accounts, and to on-premises installations.

    Keep production healthy.

    • Instrument the platform and own its monitoring and alerting.
    • Respond to incidents and debug across boundaries, from a customer"s Python traceback down to a stuck job on a GPU node.
    • Manage cloud cost and capacity as usage grows.

    Numbers & Facts

    LocationWatertown, MA

    More jobs like this

    See more jobs

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.