Spotlight Call: Thursday, September 24th at 1pm est
Location: Remote (U.S. Based)
Job Description
We are seeking a Site Reliability Engineer (SRE) to join the RADIXX team. RADIXX has built a hosted diagnostic imaging and telemedicine platform, custom built with modern web technologies and deployed on AWS in an agile environment. It gives our customers access to and collaboration on diagnostic radiology. This role owns the reliability, availability, and performance of RADIXX's production systems monitoring, incident response, automation, and infrastructure so our applications stay fast, stable, and available for customers around the clock.
What will you do?
Own the reliability, availability, and performance of RADIXX's hosted diagnostic imaging and telemedicine applications running on AWS.
Design, build, and maintain monitoring, alerting, and observability using Datadog and the AWS Console to detect and resolve issues before they affect customers.
Manage and optimize core AWS infrastructure (EC2, RDS, ECS, SQS, S3, and related services), applying infrastructure-as-code practices for repeatable, auditable deployments.
Build and maintain CI/CD pipelines in Jenkins for automated, low-risk deployment of releases and hotfixes.
Lead incident response and root cause analysis for production issues; write and maintain runbooks and postmortems to prevent recurrence.
Triage and resolve production escalations, fixing defects and shipping mini-releases with minimal customer disruption.
Define and track SLIs/SLOs and error budgets for critical services, and use that data to help prioritize engineering work.
Partner with development and product teams to build reliability, security, and scalability into new features from design through deployment.
Participate in an on-call rotation, providing timely response to production incidents.
What do you need to succeed?
3+ years of experience in site reliability engineering, DevOps, or production systems support, ideally for cloud-hosted, customer-facing applications.
Hands-on experience with AWS services (EC2, RDS, ECS, SQS, S3) and infrastructure-as-code tools (CloudFormation or Terraform).
Experience with CI/CD tooling (Jenkins or similar) and automated deployment pipelines.
Proficiency with monitoring and observability platforms (Datadog or similar) and building actionable alerting.
Scripting and automation skills (Python, Bash, or similar) to reduce manual toil.
Working knowledge of relational databases (Oracle, MySQL, SQL Server) and SQL.
Strong troubleshooting and incident-response skills, with the ability to stay calm and methodical under production pressure.
Excellent communication skills, both verbal and written, including the ability to translate technical issues to non-technical audiences.
Ability to work independently and within cross-functional teams in an agile environment.
Preferred Experience:
BA/BS in Computer Science, Engineering, or a related field, or equivalent work experience.
Experience with containerization and orchestration (Docker, ECS, or Kubernetes).
Familiarity with Java/J2EE, .NET, or similar application stacks to support root-cause debugging.
Experience handling raw image data or image files, or working in healthcare or other regulated environments.
AWS certification (SysOps Administrator, DevOps Engineer, or Solutions Architect).
Knowledge of Angular or React front-end stacks, useful for full-stack incident triage.
Let's pursue what matters together!
*** values a diverse workforce and workplace and strongly encourages women, people of color, LGBT individuals, people with disabilities, members of ethnic minorities, foreign-born residents, and veterans to apply. *** is an equal opportunity employer. Applicants will not be discriminated against because of race, color, creed, sex, sexual orientation, gender identity or expression, age, religion, national origin, citizenship status, disability, ancestry, marital status, veteran status, medical condition, or any protected category prohibited by local, state, or federal laws.