SRE Kubernetes Platform

Diverse Lynx, LLC

  • Princeton, NJ
  • 7 days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Automationunmatched
    • Budgetingunmatched
    • Cloud Architectureunmatched
    • Cloud Computingunmatched
    • CompTIA Security+unmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Customer Support/Serviceunmatched
    • DevOpsunmatched
    • Documentationunmatched
    • Federal Information Processing Standards (FIPS)unmatched
    • GitHubunmatched
    • Hybrid Cloudunmatched
    • Incident Responseunmatched
    • Jenkinsunmatched
    • Linux Operating Systemunmatched
    • Metricsunmatched
    • On Callunmatched
    • Process Improvementunmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Scripting (Scripting Languages)unmatched
    • Systems Reliabilityunmatched
    • U.S. National Institute of Standards and Technology (NIST)unmatched
    • United States Department of Defense (DoD)unmatched

    Description

    Job Description

    Must Have Technical/Functional Skills:

    • 10+ years of experience in SRE, DevOps, or infrastructure engineering
    • Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)
    • Hands-on experience working in FedRAMP High and/or DoD IL5 environments
    • Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals
    • Experience with Infrastructure as Code (Terraform preferred)
    • Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
    • Proficiency in scripting or programming (Python, Go)
    • Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)
    • Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)

    Roles & Responsibilities:

    • Design, build, and operate production-grade Kubernetes platforms in regulated environments
    • Improve system reliability through automation, thoughtful design, and continuous iteration
    • Define and drive SLOs, SLIs, and error budgets to guide reliability decisions
    • Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
    • Implement robust observability (metrics, logs, traces) to make systems understandable and actionable
    • Reduce operational toil by automating repetitive processes and improving workflows
    • Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without

    sacrificing developer velocity

    • Support ATO processes, including documentation, controls implementation, and audit readiness

    Confidential

    • Participate in on-call rotations supporting customer requests and paging alerts
    • Participate in incident response, blameless postmortems, and continuous improvement efforts
    • Help shape a platform that engineers enjoy using

    Nice to Have

    • Experience with service mesh technologies (Istio, Linkerd)
    • Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
    • Experience with GitOps workflows
    • Exposure to multi-cluster or hybrid cloud architectures
    • Knowledge of FIPS-compliant systems or DoD Cloud SRG
    • Relevant certifications (CKA, CKS, cloud provider certs, Security+)

    Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.

    Numbers & Facts

    LocationPrinceton, NJ

    Similar Jobs

    See more jobs