Site Reliability Engineer II, GovCloud

Medallia
  • McLean, Virginia
  • $103,000–$155,000 Per Year
  • Full-time
1 day ago

Job Description

Overview:

Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.  


We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees.


We empower exceptional people to create extraordinary experiences together. 


Bring your whole self.


The Role and Team

We are growing our GovCloud team and looking for a Site Reliability Engineer II to help operate and improve Medallia’s US public-sector cloud platform. You will support federal agencies and other regulated customers in a highly available, secure, and compliant environment built on AWS GovCloud and Kubernetes.

 

This is a hybrid role based near Tysons, Virginia, with regular in-office collaboration and remote flexibility. You will be a hands-on engineer who helps operate and improve production systems, works closely with engineering and security teams, and grows toward owning larger systems while helping keep the platform reliable as we scale.

Responsibilities:
  • Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services.
  • Implement, operate and optimize  AWS cloud networking, specifically managing VPCs, subnets and routing, security groups/NACLs, VPC endpoints/PrivateLink and load balancing.
  • Monitor, maintain and support production PostgreSQL — replication, backups and recovery, routine performance tuning, and upgrades — as part of the platform's data tier.
  • Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues.
  • Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and GitOps practices.
  • Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks.
  • Partner with software engineering, security, and release management teams to deploy changes safely and resolve production issues.
  • Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment.
  • Participate in an on-call rotation for production support.
  • Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries.
  • Learn the platform and grow your scope with mentorship from senior engineers.

Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.

Qualifications:

Minimum Qualifications

  • Bachelor’s degree or equivalent experience in Computer Science or a related field.
  • 2+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles (or equivalent hands-on production experience).
  • Must reside in the United States and be legally authorized to work in the US without sponsorship.
  • Production experience with:
    • Core services (IAM, compute, object storage, encryption/key management) and cloud networking on AWS (strongly preferred), Google Cloud (CGP), Azure, or a similar public cloud platform
    • Terraform or comparable infrastructure-as-code tools
    • Git and CI/CD pipelines
    • Linux and foundational systems concepts (networking, DNS, TLS/certificates)
    • PostgreSQL or another relational database — basic operations (queries, backups)
  • Familiarity with Kubernetes concepts, container orchestration, and microservices management
  • Programming and Automation: Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services.
  • Incident & Change Management: Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes.
  • Experience participating in a production on-call rotation.
  • Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems.

Preferred Qualifications

  • Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments.
  • Familiarity with security and compliance practices such as FIPS and vulnerability management.
  • Experience with observability and logging platforms in enterprise production environments.
  • Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with Redis and Kafka.
  • Experience supporting federal agencies or public-sector customers.
  • Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise.
  • Strong collaboration skills and willingness to learn in a compliance-driven environment.

 

Medallia is committed to equal pay and transparency.  The annual base salary range for this position is $103,000 - $155,000.  Please note that the salary range information provided is a general guideline and combines all of the distinct labor markets within the US.  It is uncommon for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on a variety of factors.  Medallia considers factors such as (but not limited to) scope and responsibilities of the position, candidate’s work experience, candidate’s work location, education/training, key skills, internal peer equity, external market data, as well as, market and business considerations when making compensation decisions.


Medallia also offers competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays. Benefits and eligibility may vary by location and role. 


At Medallia, we celebrate diversity and recognize the value it brings to our customers and employees. Medallia is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age (40 and over), disability, genetic information, veteran status or military service, or any other status protected by state or local law. Individuals with a disability who need an accommodation to apply please contact us at ApplicantAccessibility@medallia.com. For information regarding how Medallia collects and uses personal information, please review our Privacy Policies. Applications will be accepted for 30 days from the date this role was posted or until the role has been filled.

 

Numbers & Facts

LocationMcLean, Virginia
Job TypeFull-time
Salary$103,000–$155,000 Per Year

Skills

  • Accidental Death and Dismemberment (AD&D)unmatched
  • Amazon Web Services (AWS)unmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Change Managementunmatched
  • Cloud Computingunmatched
  • Computer Scienceunmatched
  • Computer Securityunmatched
  • Computer Skillsunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Cryptographyunmatched
  • Customer/Client Researchunmatched
  • DNS (Domain Name System)unmatched
  • Data Recoveryunmatched
  • DevOpsunmatched
  • Doctor of Nursing Science (DNS)unmatched
  • Documentationunmatched
  • Federal Governmentunmatched
  • Federal Information Processing Standards (FIPS)unmatched
  • Gitunmatched
  • GitHubunmatched
  • High Availabilityunmatched
  • Identify Issuesunmatched
  • Improvement Metricsunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Insuranceunmatched
  • Jenkinsunmatched
  • Linux Operating Systemunmatched
  • Load Balancingunmatched
  • Machine Toolunmatched
  • Maintain Complianceunmatched
  • Memory Hardwareunmatched
  • Mentoringunmatched
  • Microservicesunmatched
  • Microsoft Windows Azureunmatched
  • Network Administration/Managementunmatched
  • Network Routingunmatched
  • On Callunmatched
  • Performance Tuning/Optimizationunmatched
  • PostgreSQLunmatched
  • Privacy Controlsunmatched
  • Problem Solving Skillsunmatched
  • Process Improvementunmatched
  • Production Controlunmatched
  • Production Supportunmatched
  • Production Systemsunmatched
  • Public Cloudunmatched
  • Python Programming/Scripting Languageunmatched
  • Redisunmatched
  • Relational Databases (RDBMS)unmatched
  • Release Management/Engineeringunmatched
  • Reliability Engineeringunmatched
  • Replication and Remote Mirroringunmatched
  • Root Cause Analysisunmatched
  • SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
  • Scripting (Scripting Languages)unmatched
  • Security Patchesunmatched
  • Software Engineeringunmatched
  • Software as a Service (SaaS)unmatched
  • Subnetunmatched
  • Systems Administration/Managementunmatched
  • Team Playerunmatched
  • Technical Writingunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder