Senior Platform Engineer

Focused Labs
  • Denver, Colorado
    5 days ago

    Job Description

     

    Who we are:

    At Focused, we move quickly to deliver quality software that achieves client outcomes and meets their customer’s needs. We strategically partner with our clients to leverage our expertise in design and software, while our clients bring their own domain expertise. We work with a variety of clients from different industries, collaborating as we get new products to market, modernizing legacy systems, or helping teams learn the skills they need to be successful.   

    Our values:

    • Listen first • We are experts in product practices but life long learners in the domain of our customers. We research, collaborate, and understand. 
    • Learn why • We ask questions and talk to users to understand problem spaces, objectives, and goals, which allows us to deeply invest and drive towards the outcomes of our clients. 
    • Love your craft • We love diving into a variety of domains and solving problems.  We take pride in delivering value, in communicating progress, and guiding our clients to success.

    We are seeking an experienced Senior Platform Consultant to help organizations implement, optimize, and scale their infrastructure. 

    Key Responsibilities:

    Platform Engineering & Infrastructure

    • Augment existing infrastructure with with integrated observability solutions
    • Implement Infrastructure as Code (IaC) solutions using Terraform, Pulumi, CloudFormation, etc.
    • Architect and manage Kubernetes clusters with comprehensive monitoring and logging
    • Build CI/CD pipelines with embedded observability and automated testing

    Site Reliability Engineering (SRE)

    • Establish and maintain Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs)
    • Implement error budgets, toil reduction strategies, and capacity planning
    • Support incident response procedures and post-mortem processes

    Cloud & DevOps Engineering

    • Deploy and manage observability infrastructure across AWS, GCP, and Azure
    • Establish security, compliance, and governance frameworks for telemetry data
    • Experience automating Agent Evaluations in CI/CD pipelines and observability backends.

    Required Qualifications:

    Platform Engineering & DevOps

    • 3+ years of Platform Engineering or DevOps experience 
    • Proficiency with Infrastructure as Code tools (Terraform, Pulumi, CloudFormation, CDK)
    • Strong experience with CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, ArgoCD)

    Cloud & Infrastructure

    • Hands-on experience with major cloud providers (AWS, GCP, Azure) and their observability services
    • Experience with container technologies (Docker, Podman) and container registries
    • Knowledge of networking, security, load balancing, and distributed systems concepts

    Site Reliability Engineering

    • Experience implementing SRE practices including error budgets and toil metrics
    • Proficiency in incident management, on-call procedures, and post-mortem culture
    • Experience with capacity planning, performance optimization, and scalability design

    Programming & Automation

    • Proficiency in multiple programming languages preferred (Go, Python, Java, Node.js, Rust)
    • Strong scripting and automation skills (Bash, Python, PowerShell)
    • Understanding of software engineering best practices and testing methodologies

    Core Observability & OpenTelemetry

    • Experience in observability, monitoring, and distributed systems
    • Hands-on experience with OpenTelemetry ecosystem, including SDKs, APIs, and specifications
    • Proficiency with OpenTelemetry Collector configuration, processors, exporters, and receivers
    • Strong understanding of telemetry data models, semantic conventions, and instrumentation best practices

    Preferred Qualifications (Exceptional Candidates)

    AI & Agentic Frameworks

    • Understanding of Large Language Models (LLMs) and their application in DevOps
    • Knowledge of vector databases, embeddings, and retrieval-augmented generation (RAG)
    • Experience with AI/ML model deployment and monitoring in production environments

    Leadership & Communication

    • Strong technical writing and documentation skills
    • Ability to present complex technical concepts to diverse stakeholders
    • A passion for knowledge sharing

    Key Competencies

    • Systems thinking and ability to design holistic observability solutions
    • Strong analytical and troubleshooting skills for complex distributed systems
    • Curiosity about emerging technologies, particularly AI applications in operations
    • Adaptability to rapidly evolving cloud-native and observability technologies
    • Collaborative mindset with focus on enabling developer productivity and system reliability

    What to know before you apply: 

    • This role will require being in the Denver office three days per week and up to 20% travel within the United States.
    • Focused is unable to sponsor or take over sponsorship of the employment Visa process at this time.
    • The Denver base salary range for this role is $130,000 - $180,000.

    Numbers & Facts

    LocationDenver, Colorado

    Skills

    • Amazon Web Services (AWS)unmatched
    • Analysis Skillsunmatched
    • Application Programming Interface (API)unmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Best Practicesunmatched
    • Budgetingunmatched
    • Capacity Managementunmatched
    • Capacity and Performance Managementunmatched
    • Channel Strategiesunmatched
    • Cloud Computingunmatched
    • Consultingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Data Modelingunmatched
    • DevOpsunmatched
    • Distributed Computingunmatched
    • Divingunmatched
    • Dockerunmatched
    • Ecosystemsunmatched
    • Embedded Systemsunmatched
    • Emerging Technologyunmatched
    • GCP (Good Clinical Practices)unmatched
    • GitHubunmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Instrumentationunmatched
    • Javaunmatched
    • Jenkinsunmatched
    • Leadershipunmatched
    • Load Balancingunmatched
    • Metricsunmatched
    • Microsoft Windows Azureunmatched
    • Modeling Languagesunmatched
    • Network Securityunmatched
    • Node.jsunmatched
    • On Callunmatched
    • Performance Tuning/Optimizationunmatched
    • Presentation/Verbal Skillsunmatched
    • Problem Solving Skillsunmatched
    • Production Controlunmatched
    • Production Systemsunmatched
    • Programming Languagesunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Assurance Methodologyunmatched
    • Receiversunmatched
    • Reliability Engineeringunmatched
    • Scripting (Scripting Languages)unmatched
    • Service Level Agreement (SLA)unmatched
    • Software Designunmatched
    • Software Engineeringunmatched
    • Strategic Planningunmatched
    • Systems Reliabilityunmatched
    • Team Playerunmatched
    • Technical Presentationunmatched
    • Technical Writingunmatched
    • Telemetryunmatched
    • Test Automationunmatched
    • Windows PowerShellunmatched
    • Writing Skillsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder