Principal Engineer – Distributed Systems (GPU Edge + Inference)

Elloe AI
  • Austin, Texas
  • Remote
    30+ days ago

    Job Description

    Full-time | Remote | Infrastructure | Reports to CTO

    About Elloe
    Elloe is the trust layer for AI.
    We sit between the world’s most powerful language models and the institutions that can't afford to get it wrong — hospitals, banks, regulators. We trace and block failures in real time. That’s not marketing — we’re deployed at the European Commission, with NIH clinical trials, and inside a Top-5 EU bank catching GDPR violations live.

    This is the enforcement layer GenAI has been missing. We're not visualizing problems — we're fixing them.

    About the Role
    You’ll lead our GPU-edge inference systems. From chaos-resilient deployment to SHAP-driven compliance metrics, you’ll own global infra that makes AI safe and performant.

    What You’ll Own
    1. Global Edge Routing
    •   Design zone-routing that ensures <50ms SLA in 10+ regions
    •   Build fallback orchestration to handle compliance-aware rollbacks
    2. GPU Infra Ops
    •   Maximize utilization across 100K+ GPUs via mesh & load prediction
    •   Integrate compliance overlays with VaultChain and SHAP triggers
    3. Reliability Telemetry
    •   Ship `/vault/audit`, `/inference/predict`, `/compliance/log` endpoints
    •   Trace every edge request across governance and model layers

    Who You Are
    • Senior systems engineer with GPU fleet experience (KubeRay, Istio, Envoy)
    • Operated real-time AI infra with 10M+ QPS loads
    • Comfortable with compliance observability and infra governance

    Why This Matters
    Our competitive edge isn’t just AI — it’s defensible enforcement. This role turns that into product.

    You’ll Leave This Role With
    • Referenceable contributions to enforcement infra that’s live in EU and US institutions
    • First-hand product work across legal, engineering, and GTM teams
    • Influence over how regulatory primitives become systems people trust

    Logistics & Application
    • Start Date: Flexible (Q3–Q4 ideal)
    • Location: Remote-first; timezone overlap with NY or EU preferred
    • Compensation: Top of market salary + equity
    • To Apply: Send your resume and a sentence on the hardest infra problem you'd want to own at scale.

    Numbers & Facts

    LocationAustin, Texas (
    Remote
    )

    Skills

    • Artificial Intelligence (AI)unmatched
    • Clinical Trialunmatched
    • Distributed Computingunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Hospitalunmatched
    • Legalunmatched
    • Maintain Complianceunmatched
    • Marketingunmatched
    • Metricsunmatched
    • Modeling Languagesunmatched
    • National Institutes of Health (NIH)unmatched
    • Regulationsunmatched
    • Service Level Agreement (SLA)unmatched
    • Systems Engineeringunmatched
    • Telemetryunmatched
    • Vehicle Fleetsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder