{"id":460,"date":"2026-04-30T05:00:15","date_gmt":"2026-04-30T05:00:15","guid":{"rendered":"https:\/\/struct.ai\/articles\/improve-ai-sre-oncall-investigations\/"},"modified":"2026-04-30T05:00:15","modified_gmt":"2026-04-30T05:00:15","slug":"improve-ai-sre-oncall-investigations","status":"publish","type":"post","link":"https:\/\/struct.ai\/articles\/improve-ai-sre-oncall-investigations\/","title":{"rendered":"How to Improve AI SRE On-Call Investigations"},"content":{"rendered":"<p><em>Written by: Nimesh Chakravarthi, Co-founder &amp; CTO, Struct<\/em><\/p>\n<h2 id=\"key-takeaways\">Key Takeaways<\/h2>\n<ul>\n<li>AI SRE on-call investigations often miss silent failures like model drift, which traditional tools overlook and which waste 45+ minutes per incident on manual log hunts.<\/li>\n<li>Use a 7-step playbook with AI observability, drift alerts, and autonomous agents to cut triage time by about 80% to under 5 minutes.<\/li>\n<li>Deploy Slack-native AI agents that auto-correlate logs from Datadog, AWS, and Sentry, then generate root cause hypotheses before engineers wake up.<\/li>\n<li>Encode custom runbooks and enable conversational queries to remove context-switching and speed up MTTR while reducing engineer burnout.<\/li>\n<li><a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\">Automate your on-call runbook with Struct<\/a> for 10-minute setup, SOC2 compliance, and 85-90% helpful investigation rates at scale.<\/li>\n<\/ul>\n<h2>Pinpoint Your Current AI SRE On-Call Gaps<\/h2>\n<p>Start by auditing your current on-call reality. Track weekly triage hours, false positive rates, and escalation frequency. Most teams discover they spend 60-80% of incident time on manual investigation, which becomes the exact bottleneck AI agents remove.<\/p>\n<p>AI systems introduce failure modes that traditional SRE practices rarely catch. A 2025 McKinsey Global AI survey found that 51% of organizations using AI experienced at least one negative consequence from AI risks, and nearly one-third reported issues from AI inaccuracy. Your current Datadog and GCP monitoring likely catches infrastructure issues but misses semantic drift, hallucinations, and embedding anomalies.<\/p>\n<p>To identify these gaps in your own environment, create this checklist to quantify your baseline: alert volume per week, average MTTR for P1 incidents, tools that require manual correlation such as Datadog, Sentry, and AWS, and gaps in AI-specific monitoring like concept drift detection. These metrics reveal whether your team operates reactively or proactively. Most teams find they wait for user complaints instead of catching degradation early.<\/p>\n<p><a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\">Connect Struct to Your Slack for Instant Auto-Triage<\/a><\/p>\n<h2>7 Steps to Automate AI SRE On-Call Investigations<\/h2>\n<h3>Step 1: Deploy an AI-Focused Observability Stack<\/h3>\n<p>Layer AI-specific monitoring on top of your existing Datadog and AWS infrastructure. Track stability such as successful model responses versus failures, latency, load, model drift, data drift, and cost metrics like token usage. This creates a reliable data foundation for automated investigations.<\/p>\n<h3>Step 2: Configure Targeted Drift Detection Alerts<\/h3>\n<p>Set up Prometheus-style alerts for AI-specific failures. <a href=\"https:\/\/agility-at-scale.com\/ai\/generative\/continuous-evaluation-and-drift-monitoring\" target=\"_blank\" rel=\"noindex nofollow\">Use Population Stability Index (PSI) thresholds, where below 0.1 indicates negligible drift, 0.1 to 0.25 signals moderate drift that needs investigation, and above 0.25 shows significant distributional change that requires immediate action<\/a>. These thresholds give your team clear triggers instead of vague warnings.<\/p>\n<h3>Step 3: Deploy an AI SRE Agent in Slack and PagerDuty<\/h3>\n<p>Connect an AI agent directly to your Slack and PagerDuty channels. When alerts fire, the agent pulls logs, correlates traces, and generates root cause hypotheses in under 5 minutes. This work happens before you even open your laptop.<\/p>\n<h3>Step 4: Encode Your Team\u2019s Custom Runbooks<\/h3>\n<p>Turn your team\u2019s tribal knowledge into executable runbooks. Capture correlation IDs, specific log patterns, and clear escalation procedures. Auto-triggered diagnostic runbooks save 15-30 minutes of investigation time by collecting service health checks and log analysis before engineers acknowledge alerts. This structure makes every incident feel guided instead of ad hoc.<\/p>\n<h3>Step 5: Generate Dynamic Incident Dashboards<\/h3>\n<p>Once your runbooks execute automatically, you need clear visibility into what they uncover. Create incident-specific dashboards that merge Datadog metrics, AWS traces, and Sentry exceptions into unified timelines. This approach removes context-switching between several tools during outages and keeps everyone aligned on a single view.<\/p>\n<h3>Step 6: Enable Conversational AI for Live Triage<\/h3>\n<p>Deploy Slack-native AI that responds to natural language queries such as \u201cpull logs from 5 minutes prior\u201d or \u201ccheck if this impacts user segment X.\u201d This capability replaces manual log hunting during active incidents and keeps engineers focused on decisions instead of data gathering.<\/p>\n<h3>Step 7: Automate Handoffs to Code and Postmortems<\/h3>\n<p>After the team confirms root cause, automatically generate GitHub pull requests with suggested fixes or hand off context to coding agents for implementation. This automation shortens the gap between diagnosis and remediation and keeps follow-through consistent.<\/p>\n<table>\n<tr>\n<th>Approach<\/th>\n<th>Setup Time<\/th>\n<th>Accuracy<\/th>\n<th>Triage Reduction<\/th>\n<\/tr>\n<tr>\n<td>Open-Source Agents<\/td>\n<td>Days<\/td>\n<td>Varies<\/td>\n<td>Varies<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">Struct<\/a><\/td>\n<td>10 minutes<\/td>\n<td>Varies<\/td>\n<td>Varies<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/www.nature.com\/articles\/srep17168\/tables\/3\" target=\"_blank\" rel=\"noindex nofollow\">Manual Investigation<\/a><\/td>\n<td>42 hrs<\/td>\n<td>91.6%<\/td>\n<td>N\/A<\/td>\n<\/tr>\n<\/table>\n<p><a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\">See How Struct Executes These 7 Steps Automatically<\/a><\/p>\n<h2>Wire AI Agents into Your Existing Incident Workflow<\/h2>\n<p>Effective AI SRE ties directly into the tools your team already uses. PagerDuty alerts trigger Slack notifications, which activate AI agents that query Datadog and Sentry, then post findings back to Slack with GitHub pull request links for fixes. This flow keeps the entire incident lifecycle inside your communication hub.<\/p>\n<p>Struct\u2019s edge comes from being Slack-native, so you avoid tool switching during 3 AM incidents. <a href=\"https:\/\/datadoghq.com\/blog\/bits-ai-sre-deeper-reasoning\" target=\"_blank\" rel=\"noindex nofollow\">Datadog\u2019s Bits AI SRE supports direct triage actions such as sending Slack messages, creating incidents, and generating Jira tickets with prefilled context to reduce context switching<\/a>. Struct extends this pattern across your full stack and runbooks.<\/p>\n<p>Design unified handoff workflows where AI investigations automatically populate post-mortem templates and create follow-up tickets. This approach closes the loop from detection to resolution while keeping manual overhead low.<\/p>\n<h2>Measure AI SRE Impact and Improve Over Time<\/h2>\n<p>Track four key metrics: triage time reduction with a target of about 80%, MTTR improvement, investigation accuracy rate, and engineer satisfaction scores. <a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">Struct\u2019s fintech customers achieved the 80% reduction mentioned earlier, with real-world validation across multiple incident types<\/a>.<\/p>\n<p>Set weekly review cadences and run chaos engineering simulations to test your automated workflows. <a href=\"https:\/\/ai2roi.substack.com\/p\/ai-to-roi-metric-mean-time-to-resolve\" target=\"_blank\" rel=\"noindex nofollow\">Mature implementations achieve 75-82% MTTR reductions over 12-20 months through iterative improvements<\/a>. These reviews keep your AI agents aligned with evolving systems and failure modes.<\/p>\n<h2>Common Pitfalls, Proven Practices, and AI SRE Tool Comparison<\/h2>\n<p>Avoid reactive AI usage such as manually prompting ChatGPT during incidents. That pattern hits context limits, depends on copy-paste workflows, and needs constant human guidance. Proactive agents instead start investigations automatically when alerts fire and arrive with context already assembled.<\/p>\n<p>Follow a few best practices to keep your rollout safe and effective. Ensure SOC2 compliance for sensitive data, start junior engineers with AI-assisted investigations for faster onboarding, and maintain a composable architecture that adapts to your specific tech stack.<\/p>\n<table>\n<tr>\n<th>Tool<\/th>\n<th>Setup Time<\/th>\n<th>Accuracy<\/th>\n<th>Triage Reduction<\/th>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">Struct<\/a><\/td>\n<td>10 minutes<\/td>\n<td>Varies<\/td>\n<td>Varies<\/td>\n<\/tr>\n<tr>\n<td>Cleric.ai<\/td>\n<td>Weeks<\/td>\n<td>Varies<\/td>\n<td>Varies<\/td>\n<\/tr>\n<tr>\n<td>Claude\/ChatGPT<\/td>\n<td>Manual<\/td>\n<td>Varies<\/td>\n<td>Varies<\/td>\n<\/tr>\n<tr>\n<td>Open-source<\/td>\n<td>Days<\/td>\n<td>Varies<\/td>\n<td>Varies<\/td>\n<\/tr>\n<\/table>\n<h2>Accelerate AI SRE On-Call with Struct<\/h2>\n<p>These 7 steps turn reactive fire-fighting into proactive AI-powered investigations. <a href=\"https:\/\/aws.amazon.com\/blogs\/devops\/leverage-agentic-ai-for-autonomous-incident-response-with-aws-devops-agent\" target=\"_blank\" rel=\"noindex nofollow\">AWS DevOps Agent customers report faster investigations and lower MTTR<\/a>. Autonomous agents now handle the heavy lifting while your team focuses on high-value decisions.<\/p>\n<p><a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\">Stop 3AM Log Hunts and Set Up Struct in 10 Minutes<\/a><\/p>\n<h2>FAQ<\/h2>\n<h3>What is an AI SRE agent and how does it differ from open-source options?<\/h3>\n<p>An AI SRE agent automatically investigates alerts by pulling logs, correlating traces, and generating root cause hypotheses without human intervention. Unlike open-source tools that require days of setup and manual configuration, Struct provides a proactive platform that deploys in 10 minutes and learns your specific architecture patterns for increasingly accurate investigations over time.<\/p>\n<h3>How long does it take to set up automated AI SRE investigations?<\/h3>\n<p>Struct connects to your existing Slack, GitHub, and observability tools in under 10 minutes. You authenticate your integrations, configure alert channels, and then start receiving automated investigation reports immediately. You avoid lengthy enterprise deployments and complex indexing.<\/p>\n<h3>Is this compliant with HIPAA and SOC2 requirements?<\/h3>\n<p>Yes, Struct is fully <a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">SOC2 Type II and HIPAA compliant<\/a>. Your logs are processed ephemerally without persistent storage, which meets strict enterprise security requirements while still enabling automated investigations.<\/p>\n<h3>What if our logging and telemetry infrastructure needs improvement?<\/h3>\n<p>AI agents need basic observability foundations such as Datadog metrics, structured logs with correlation IDs, and alert triggers. If your system lacks fundamental logging or trace IDs, focus on improving data quality first. Struct works best with teams that already use modern observability platforms.<\/p>\n<h3>Can junior engineers safely handle on-call with AI assistance?<\/h3>\n<p>Yes. AI agents provide strong starting points for every alert, including context, timeline, and suggested next steps. This support allows newer engineers to take on-call shifts confidently while they learn your system architecture through guided investigations.<\/p>\n<h3>How does this compare to Datadog\u2019s built-in AI features?<\/h3>\n<p>Datadog\u2019s Bits AI provides investigation capabilities inside their platform. Struct adds seamless Slack integration, custom runbook encoding, and automated handoffs to GitHub for end-to-end incident resolution without leaving your communication hub. Together, these features create a more complete incident workflow.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Cut AI SRE triage time by 80% with autonomous agents &amp; drift alerts. Struct&#8217;s 7-step playbook reduces MTTR to under 5 minutes. Start free.<\/p>\n","protected":false},"author":73,"featured_media":459,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"inline_featured_image":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-460","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/posts\/460","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/comments?post=460"}],"version-history":[{"count":0,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/posts\/460\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/media\/459"}],"wp:attachment":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/media?parent=460"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/categories?post=460"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/tags?post=460"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}