Sunrise Systems Inc logo

Software Development Engineer 5

Sunrise Systems Inc

  • San Jose, CA
  • 7 days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Analysis Skillsunmatched
    • Automationunmatched
    • Budgetingunmatched
    • Cadenceunmatched
    • Code Reviewsunmatched
    • Computer Scienceunmatched
    • Debugging Skillsunmatched
    • Device Driversunmatched
    • Dockerunmatched
    • Engineeringunmatched
    • Establish Prioritiesunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Graphicsunmatched
    • JSONunmatched
    • Kernel Programmingunmatched
    • Machine Learningunmatched
    • Machine Toolunmatched
    • Memory Hardwareunmatched
    • Metricsunmatched
    • Operational Support Systems (OSS)unmatched
    • Performance Analysisunmatched
    • Performance Engineeringunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Regression Testingunmatched
    • Software Developmentunmatched
    • Software Engineeringunmatched
    • Technical Leadershipunmatched
    • Technical Writingunmatched
    • Telemetryunmatched
    • Test Suiteunmatched
    • Testingunmatched
    • Value Engineeringunmatched

    Description

    Position Overview

    • Job Title: Software Development Engineer 5 - Agentic ML Infrastructure
    • Job Class: Engineering
    • Level: SDE 5
    • Engagement: 6-month contract, full-time (1.0 FTE)
    • Location: United States - remote/hybrid per AMD policy
    • Reports To: Technical Lead, AMD AGI

    About the Role

    AMD is building agentic software to improve how teams develop, deploy, and operate machine learning workloads on AMD GPUs.

    This senior contract role focuses on designing and delivering production-grade intelligent agents across the AMD software stack, including:

    • Training and inference frameworks
    • Cluster tooling
    • Performance analysis
    • Operational workflows for large-scale GPU deployments

    The role will lead implementation of Python systems that combine LLM orchestration, tool integration, and rigorous evaluation to automate high-value engineering work, including:

    • Investigating distributed job failures
    • Interpreting telemetry and profilers
    • Accelerating performance analysis
    • Reducing manual toil in ML operations

    This is a senior individual contributor engagement. The engineer is expected to own significant components end-to-end, make sound technical decisions within program direction, and deliver maintainable code with minimal supervision.

    Key Responsibilities

    Agent Framework & LLM Engineering

    • Architect and implement Python agent frameworks, including:

    • Tool adapters

    • Orchestration loops

    • Structured outputs

    • Session handling

    • CLI or service interfaces suitable for production use

    • Build LLM-powered workflows using LiteLLM, DSPy, or equivalent.

    • Implement bounded prompts, allowlisted tools, budget caps, structured JSON outputs, and evidence-backed reporting.

    ML & GPU Infrastructure

    • Integrate agents with the AMD GPU and ML software ecosystem, including:

    • ROCm tooling

    • PyTorch

    • Distributed training and inference stacks

    • Profilers

    • Cluster metrics

    • Log pipelines

    • Debug utilities

    • Solve ML operations problems on training and inference clusters, including:

    • Multi-node failure analysis

    • Performance regression investigation

    • Configuration and launch issues

    • Operational automation

    Quality, Testing & Evaluation

    • Own agent quality and evaluation infrastructure.
    • Develop fixture-driven tests and regression suites.
    • Implement JSON Schema validation and accuracy metrics.
    • Integrate testing and evaluation into CI before rollout.

    Deployment & Documentation

    • Deliver deployment-ready artifacts, including:

    • Docker packaging

    • Kubernetes packaging

    • Runbooks

    • Technical documentation

    • Apply security-conscious defaults for customer-controlled environments.

    Collaboration & Delivery

    • Partner with framework, performance, and field teams.
    • Incorporate code review feedback and iterate based on production and pilot learnings.
    • Own significant technical components independently and deliver production-quality solutions with minimal supervision.

    Required Qualifications

    Education & Experience

    • Bachelor''s degree or higher in:

    • Computer Science

    • Computer Engineering

    • Electrical Engineering

    • Related field

    • Equivalent practical experience accepted.

    • 10+ years of professional software engineering experience.

    • 5+ years of experience developing production Python systems at scale.

    Distributed ML & GPU Systems

    • Deep experience with distributed ML training or inference on GPU clusters.

    • Experience with:

    • Multi-node jobs

    • Collective communication failures

    • Log and metric correlation across ranks and nodes

    • Strong hands-on PyTorch experience.

    • Production familiarity with large-scale ML stacks such as:

    • Megatron-LM

    • DeepSpeed

    • TorchTitan

    • vLLM

    • Equivalent technologies

    ML Debugging & Performance Engineering

    • Proven ML debugging and performance engineering experience across multiple failure modes, including:
    • Throughput regression
    • Memory errors
    • Numerical instability
    • Misconfiguration
    • Profiler or trace analysis

    LLM Agents & Software Engineering

    • Demonstrated experience delivering LLM agent or orchestration systems in production or near-production environments.

    • Experience with:

    • Tool routing

    • Structured outputs

    • Reliability

    • Testability

    • Track record of delivering complex software on schedule.

    • Strong capabilities in:

    • Clean, maintainable code

    • Automated testing

    • Code review

    • Technical communication

    • Ability to work independently, prioritize ambiguous requirements, and align weekly with a technical lead.

    Preferred Qualifications

    • Open-source ML development, including:

    • Upstream framework repositories

    • Community CI patterns

    • OSS tooling integration

    • ML performance optimization, including:

    • Parallelism

    • Throughput/MFU tuning

    • Roofline analysis

    • Profiler-driven investigation

    • ROCm and GPU cluster operations, including:

    • Metrics

    • Profilers

    • Health monitoring

    • Slurm

    • Kubernetes job environments

    • Security-aware agent deployment, including:

    • Secret handling

    • Egress control

    • Customer VPC constraints

    • On-premises deployment constraints

    Working Model

    • Duration: 12 months
    • Schedule: Full-time contract
    • Cadence: Weekly sync with technical lead
    • Delivery Model: Milestone-driven
    • Scope: Implementation across AMD agent and ML infrastructure programs
    • Priorities: Allocation may shift based on program priorities at the direction of the technical lead

    What This Role Is Not

    This position is not focused on:

    • Graphics driver development
    • Kernel development
    • Low-level GPU driver development
    • Research-only work without production deliverables
    • Junior or mid-level implementation

    SDE 5 expectations: Senior ownership, strong technical depth, and demonstrated expertise across ML systems and agent engineering.

    Numbers & Facts

    LocationSan Jose, CA
    IndustryStaffing/Employment Agencies
    Company Size100 to 499 employees
    Year Founded1990
    Websitehttp://www.sunrisesys.com/

    About Company

    Sunrise Systems was founded in 1990 with a clear vision to deliver world-class staffing service solutions in all labor categories, including IT consulting and solutions; all with the commitment to provide service that exceeds expectations and become the most trusted name in the industry. More than two and a half decades later, we pride ourselves on being at the forefront of the staffing industry. Combining our deep industry expertise, insights, and global resources, we have partnered with our clients to connect them with top professionals across several different industries.

    We provide cost-effective Managed Staffing Solutions, Information Technology and Information Technology Consulting Services to several Fortune 500 companies and U.S. Government agencies. We provide our clients with flexible engagement models and customized products that are budget and time specific. Understanding the challenges that every business faces, we offer our services either on-site at the clients' site or from one of our globally distributed technology centers. Our onshore and offshore development capabilities ensure that we excel at meeting customer requirements every single time.

    Our collective business experience spans over two and a half decades and ranges from:

    • Business, management, and technical fields
    • Information technology consulting and software solutions.
    • Providing strategic support for the development and long-term growth of new business ventures across several industries including but not limited to; accounting, banking, finance, and recruitment.
    • Motivating technology staff and establishing partnerships with Fortune 500 companies

    Sunrise Systems has a vast range of competence in:

    • Design, development, and support of cloud-based solutions from simple to highly complexed
    • Database administration of multi-platform applications, complex databases, and web-based environments that include all aspects of installation, planning, maintenance, and monitoring.
    • Data processing and data migration
    • Application re-engineering and platform migration
    • Working with the Information Systems and end-user communities at all levels to resolve issues and establish consensus.

    Similar Jobs

    See more jobs