Technical Support Principal Engineer - AI Network

Upscale AI Inc

NY

JOB DETAILS
SALARY
$248,000–$269,000 Per Year
SKILLS
ASIC (Application Specific Integrated Circuit), Ansible, Application Programming Interface (API), Artificial Intelligence (AI), Automation, Automation Systems, Benchmarking, Code Reviews, Communication Skills, Computer Engineering, Computer Programming, Computer Science, Cross-Functional, Customer Experience, Customer Relations, DOM (Document Object Model), Debugging Skills, Debugging Tools, Docker, Documentation, Electrical Engineering, Ethernet, Funding, GPU (Graphics Processing Unit), Go Programming Language (Golang), Hardware Architecture, Hardware Quality Assurance, IP (Internet Protocol) Routing, Identify Issues, Integration Testing, JSON, Linux Operating System, Machine Tool, NFS (Network File System), National Intelligence Council (NIC), Network Administration/Management, Network Operating Systems, Network Operations Center, Network Performance/Analysis, Network Routing, Network Switching, Network Systems, Operating Systems, Optics, PHY, Performance Analysis, Performance Metrics, Presentation/Verbal Skills, Problem Solving Skills, Proof of Concept, Python Programming/Scripting Language, QoS (Quality of Service), Quality Assurance, REST (Representational State Transfer), Root Cause Analysis, Signal Integrity, Startup, Strategic Planning, Switching Architecture, TCP/IP (Transmission Control Protocol/Internet Protocol), Team Player, Technical Support, Technical Writing, Telemetry, Test Automation, VLAN (Virtual Local Area Network), Writing Skills
LOCATION
NY
POSTED
2 days ago

Technical Support Principal Engineer - AI Network

US - Headquarters

Support /

Reg - Full-Time /

On-site

apply for this job

Why join Upscale AI

Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world's most demanding AI workloads.

We focus on first-principles engineering across silicon, systems, and networking-where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.

If you're looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI-Upscale AI is where you can produce meaningful work at the frontier-and operate at a high standard.

Role Overview

We are looking for a Technical Support Lead Engineer with 12+ years of experience who is passionate about building reliable networking switches and solving challenging infrastructure problems.

Own end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You'll be the leader for switching ASIC pipeline tuning, platform readiness, and fabric performance (Ethernet/RoCEv2 and/or InfiniBand), turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.

As part of the AI Networking team, you''ll debug software, develop automation, validate networking platforms, troubleshoot complex system issues, and work closely with engineering teams and customers to ensure the best customer experience. You''ll have the opportunity to work on technologies powering next-generation GPU clusters and AI data centers while learning from experienced engineers in a fast-paced startup environment.

Responsibilities

Design, develop, and maintain features and enhancements for the SONiC NOS platform.

  • Debug, troubleshoot, and resolve issues on SONiC platforms.
  • Develop and execute debugging and troubleshoot infrastructure.
  • Collaborate closely with cross-functional teams including hardware engineers and Test teams.
  • Participate in code reviews, architecture discussions, and documentation efforts.
  • Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
  • Develop debugging tools, for Traffic monitoring and Performance measurements.
  • Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
  • Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
  • Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
  • Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
  • Contribute to backend services, APIs, orchestration components, and infrastructure automation.
  • Document technical findings and communicate effectively with engineering teams and customers.

Qualifications

Required Qualifications

  • Bachelor's or master's degree in computer science, Electrical Engineering, or a related field.
  • Minimum of 12 years of work experience is required, with at least 3 years of hands-on SONiC or equivalent Network Operating System (NOS) development experience preferred.
  • Strong programming skills in Python, Go, or a similar language.
  • Solid understanding of Linux, TCP/IP networking, routing, switching, VLANs, and network troubleshooting.
  • Experience with PTF (Packet Test Framework) and SPyTest for network validation.
  • Familiarity with Linux internals, docker containers.
  • Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
  • Knowledge of network ASICs and switch hardware architecture is mandatory.
  • Excellent written and verbal communication skills.
  • Ability to thrive in a collaborative, fast-paced startup environment with a strong sense of ownership.

Preferred Qualifications

  • Data-center networking with hands-on switch ASIC tuning and platform bring-up.
  • Proven RoCEv2 deployments at 200/400/800 G: ECN/PFC design, DCQCN tuning, DSCP/PCP mapping, queue/WRR shaping.
  • Deep buffer/queueing knowledge (headroom, xon/xoff, dynamic thresholds, VOQ vs shared pools).
  • SONiC (buffers.json, qos.json, pfcwd, ecn), Enterprise-OS (class-map/type qos & network-qos, policy-map, queuing).
  • Optics & PHY: PAM4 signal integrity, FEC modes, DOM/RS-FEC counters, AN/LT quirks, DAC/AOC selection.
  • Tooling: ethtool, devlink, mlnx_qos, perfquery, switch telemetry (INT/sFlow/ERSPAN), gNMI/REST, Prometheus.
  • Benchmarking: perftest, nccl-tests, iPerf3, traffic generators; reading queue stats, ECN marks, and CNP behavior.
  • EVPN/VXLAN leaf-spine for AI pods; flowlet or latency-aware hashing.
  • BlueField DPU offloads, GPUDirect RDMA; NIC QoS (DSCP-to-TC, PFCx, GEARBOX/FW).
  • Storage fabrics for AI (NFS-RDMA, NVMe-oF) and their QoS interactions.
  • Python/Ansible for templated ASIC profiles; Git workflows for config promotion.

$248,000 - $269,000 a year

Where you fall within that range depends on your experience, skills, and impact-we benchmark against internal levels to keep things fair and consistent.

Equal Opportunity

Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We're proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.

Accessibility & Accommodations

We're committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at [email protected]-we're happy to help. Note: This inbox is only for accommodation requests.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

apply for this job

About the Company

U

Upscale AI Inc