Extensive experience with Site Reliability Engineering (SRE) principles including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering practices Proven expertise with Azure cloud services including compute, storage, networking, monitoring, and platform-as-a-service (PaaS) offerings with deep understanding of operational best practices Strong experience with monitoring and observability tools (Azure Monitor, Application Insights, Grafana, Prometheus, ELK stack) and implementing alerting, dashboards, and log aggregation Demonstrated ability to troubleshoot complex technical issues across application, platform, and infrastructure layers with strong analytical and problem-solving skills Desired – Experience operating AI/ML platforms, large language model services (Azure OpenAI), data analytics platforms (Databricks, Synapse), or high-scale cloud applications in production environments Hands-on experience with Azure Government or other secure government cloud environments (AWS GovCloud) with understanding of compliance monitoring, security operations, and federal operational requirements Background in federal government, mission-critical systems, or 24/7 operational environments with experience supporting incident response, change management, and operational excellence programs. Job Title: SRE Platform EngineerJob Category: Information TechnologyTime Type: Full timeMinimum Clearance Required to Start: NoneEmployee Type: RegularPercentage of Travel Required: NoneType of Travel: None* * * The Opportunity: CACI is seeking a seasoned Site Reliability (SRE) Platform Engineer to support the Department of Homeland Security (DHS) Office of the Inspector General (OIG).