Monitor Dynatrace consumption, identify unexpected telemetry growth, support cost attribution, and recommend optimizations that preserve required monitoring coverage.
Deploy and support OneAgent, Smartscape, Synthetics, and Real User Monitoring (RUM)
Build dashboards and KPIs aligned to business outcomes
Maintain monitoring standards, event management, knowledge articles, and operational runbooks.
Perform analytics and reporting on the Dynatrace platform
Support platform upgrades and future adoption
Manage vendor escalations with Dynatrace support
Collaborate with internal teams to customize Dynatrace solutions to meet specific application and infrastructure needs
Support onboarding of new applications into Dynatrace
Validate Monitoring requirements and acceptance criteria
Assist application teams with instrumentation and monitoring design.
Verify ServiceNow routing and alert configuration prior to operational acceptance.
Review and improve alert quality by addressing duplicate, noisy, unactionable, incorrectly scoped, or incorrectly routed Dynatrace Problems and notifications
Drive application performance monitoring (APM)
Perform anomaly detection and root cause analysis assisting technology teams
Correlate performance data across application and infrastructure layers
Support approved integrations between Dynatrace and infrastructure, cloud, network, and ITSM monitoring
Help integrate Dynatrace with network and infrastructure monitoring tools
Collaborate with teams using network performance and flow tools
Correlate telemetry across application, network, and infrastructure layers
Administer and support approved Dynatrace monitoring for Azure-hosted applications, infrastructure, and platform services
Use scripting or automation to improve repeatability, configuration quality, reporting, and support efficiency
Collaborate with cloud and network teams on connectivity, access, and platform prerequisites
Diagnose Dynatrace product, collection, configuration, integration, and access issues.
Investigate OneAgent, ActiveGate, extension, synthetic-monitoring, and telemetry-collection failures.
Validate telemetry availability and freshness during incidents.
Use metrics, logs, traces, topology, events, and synthetic results to support technical investigations.
Partner with application, infrastructure, network, cloud, and ServiceNow teams to route issues to the team capable of restoration.
Document investigations, corrective actions, known gaps, and vendor escalations in the applicable ITSM records.
Participate in problem management and continuous-improvement activities for recurring failures, alert noise, routing issues, and monitoring blind spots.