{"id":831,"date":"2026-08-14T05:01:56","date_gmt":"2026-08-14T05:01:56","guid":{"rendered":"https:\/\/struct.ai\/articles\/automated-troubleshooting-ai\/"},"modified":"2026-08-14T05:01:56","modified_gmt":"2026-08-14T05:01:56","slug":"automated-troubleshooting-ai","status":"publish","type":"post","link":"https:\/\/struct.ai\/articles\/automated-troubleshooting-ai\/","title":{"rendered":"Automated Troubleshooting with AI: Cut MTTR Fast"},"content":{"rendered":"<p><em>Written by: Nimesh Chakravarthi, Co-founder &amp; CTO, Struct<\/em><\/p>\n<h2 id=\"key-takeaways\">Key Takeaways<\/h2>\n<ul>\n<li>Automated troubleshooting with AI ingests alerts from Datadog and Sentry, runs regression analysis across metrics, logs, and traces, then surfaces only actionable anomalies within minutes.<\/li>\n<li>Struct correlates Sentry issues with Datadog metrics, cloud infrastructure, GitHub deploy history, and logs to deliver cited root-cause hypotheses in under 5 minutes.<\/li>\n<li>After the root cause is confirmed, Struct hands off context to a local CLI, an AI coding agent, or generates a pull request while respecting production guardrails.<\/li>\n<li>Incident resolution verification re-queries observability data every minute to confirm fixes and automatically updates Slack status without manual dashboard checks.<\/li>\n<li>Struct automates your on-call runbook to cut triage time 80% in 10 minutes\u2014<a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\"><strong>start your 30-day risk-free pilot<\/strong><\/a>.<\/li>\n<\/ul>\n<p>Before diving into Struct\u2019s workflow, here is how it compares to the AI features built into Datadog and Sentry. The key difference is that Struct correlates signals across both tools, while Datadog Bits AI and Sentry Seer each remain scoped to their own telemetry.<\/p>\n<h2>Tool Comparison: Datadog Bits AI vs. Sentry Seer vs. Struct<\/h2>\n<table>\n<thead>\n<tr>\n<th>Tool<\/th>\n<th>Pricing<\/th>\n<th>Key Integrations<\/th>\n<th>Limitation<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Datadog Bits AI<\/td>\n<td><a href=\"https:\/\/www.nobs.tech\/blog\/datadog-bits-ai-pricing-ai-credits-governance\" target=\"_blank\" rel=\"noindex nofollow\">Datadog Bits AI is billed separately via consumption-based AI Credits (e.g., ~$500 for 500 credits on annual commit), not bundled into the core platform.<\/a><\/td>\n<td>Datadog metrics, logs, APM, dashboards<\/td>\n<td><a href=\"https:\/\/struct.ai\/blog\/struct-vs-datadog\" target=\"_blank\">Scoped to Datadog telemetry, with limited cross-stack correlation with Sentry or GitHub deploy history<\/a><\/td>\n<\/tr>\n<tr>\n<td>Sentry Seer<\/td>\n<td><a href=\"https:\/\/docs.sentry.io\/pricing\/quotas\/manage-seer-budget\/\" target=\"_blank\" rel=\"noindex nofollow\">Sentry Seer is available as a paid add-on to Team or Business plans at $40 per active contributor per month.<\/a><\/td>\n<td>Sentry issues, stack traces, releases<\/td>\n<td><a href=\"https:\/\/struct.ai\/blog\/struct-vs-sentry-seer\" target=\"_blank\">Limited to Sentry telemetry, and does not ingest Datadog metrics, cloud logs, or infrastructure signals<\/a><\/td>\n<\/tr>\n<tr>\n<td>Struct<\/td>\n<td>Startup (30 issues\/mo, up to 5 users), Growth (200 issues\/mo, unlimited users), Enterprise (custom); 30-day risk-free pilot included<\/td>\n<td><a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">Slack, Datadog, Sentry, GitHub, AWS CloudWatch, GCP Logs, Azure, Grafana, PagerDuty, Linear, Jira<\/a><\/td>\n<td>Requires access to logs via integrations, and does not currently support full on-premise (zero-egress VPC) deployment<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Step 1: Anomaly Detection That Cuts Noise<\/h2>\n<p>Struct ingests alerts from Datadog and Sentry the moment they fire, runs regression analysis and correlation across metrics, logs, and traces, then surfaces only actionable anomalies within minutes. <a href=\"https:\/\/blog.opsramp.com\/it-alerts-aiops-savings\" target=\"_blank\" rel=\"noindex nofollow\">Industry AIOps implementations can reduce alert volume by more than 90%<\/a>, showing how powerful automated filtering can be when applied correctly. Struct\u2019s intelligent deduplication filters transient noise before any engineer is paged, so the alerts that reach your team already carry confirmed signal.<\/p>\n<p><a href=\"https:\/\/devops.com\/the-end-of-alert-fatigue-how-ai-powered-observability-is-transforming-sre-teams-in-2026\" target=\"_blank\" rel=\"noindex nofollow\">The Catchpoint SRE Report 2025 found that nearly 70% of SREs report on-call stress has contributed to burnout and attrition.<\/a> Struct tackles that problem at the source by eliminating false-positive pages at the detection stage.<\/p>\n<h2>Step 2: Root-Cause Analysis in Minutes<\/h2>\n<p><a href=\"https:\/\/struct.ai\/blog\/struct-vs-sentry-seer\" target=\"_blank\">Struct ingests Sentry issues the moment they fire and correlates them with Datadog metrics, cloud infrastructure, GitHub deploy history, logs, traces, and other tools to produce a cited root-cause hypothesis.<\/a> The system maps exceptions, recent commits, and infrastructure signals into a single unified timeline. Engineers no longer need to pivot manually between tools to reconstruct what happened.<\/p>\n<p>The results are concrete. <a href=\"https:\/\/struct.ai\/case-study\/arcana\" target=\"_blank\">Arcana, a Series A fintech with over 40 engineers, reduced median investigation time from 30 minutes to 2 minutes and reclaimed 56 developer hours per month after integrating Struct with Sentry, GitHub, GCP Cloud Logging, and Slack.<\/a> <a href=\"https:\/\/struct.ai\/blog\/struct-vs-datadog\" target=\"_blank\">Senior engineer hours spent on investigation dropped from approximately 60 to 4 per month.<\/a><\/p>\n<p><a href=\"https:\/\/stackgen.com\/blog\/how-to-automate-alert-triage-with-ai-sres\" target=\"_blank\" rel=\"noindex nofollow\">Across the industry, 60% or more of incident MTTR is consumed not in remediation but in diagnosis, including establishing what is broken, assembling the right people, and reconstructing the chain of causation.<\/a> Struct compresses that diagnostic phase to under 5 minutes.<\/p>\n<h2>Step 3: Automated Remediation With Guardrails<\/h2>\n<p>Once the root cause is confirmed, Struct hands off full context to a local CLI, an AI coding agent, or directly generates a pull request while respecting production guardrails. <a href=\"https:\/\/aws.amazon.com\/blogs\/devops\/leverage-agentic-ai-for-autonomous-incident-response-with-aws-devops-agent\" target=\"_blank\" rel=\"noindex nofollow\">Preview customers using agentic AI for autonomous incident response reported up to 75% lower MTTR, 80% faster investigations, and 94% root-cause accuracy.<\/a> Those kinds of gains depend on a careful balance between automation and human control.<\/p>\n<p>Struct follows the governed-autonomy model that has become the 2026 standard. Agents propose changes, engineers authorize them, and every step is recorded. The handoff to a coding agent or pull request remains a bounded action, not a unilateral production change.<\/p>\n<h2>Step 4: Incident Resolution Verification in Slack<\/h2>\n<p>Incident resolution verification closes the loop by re-querying observability data to confirm an issue has returned to the expected state. Self-healing ITOps treats validation as a first-class stage that confirms system performance has returned to expected levels using observability signals, and runs validation checks before and after executing approved actions. If the issue remains unresolved, it rolls back the change and escalates with details on what was attempted.<\/p>\n<p>Struct\u2019s Incident Tracker runs this verification loop approximately every minute and updates incident status automatically in Slack. No engineer needs to manually re-check dashboards to confirm a fix held. <a href=\"https:\/\/struct.ai\/case-study\/arcana\" target=\"_blank\">Arcana now runs 2,100+ automated investigations monthly with an 85%\u201390%+ helpful rate<\/a>, and resolution verification closes the loop on each one.<\/p>\n<p><a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\"><strong>See Struct\u2019s incident verification in action\u2014book a demo<\/strong><\/a><\/p>\n<h2>Step 5: Slack-Native Collaboration for Incidents<\/h2>\n<p>All investigation output, timelines, and suggested fixes appear directly in the original Slack alert thread. Engineers can ask follow-up questions, request additional logs, or test alternative hypotheses without leaving the channel. Struct\u2019s conversational interface accepts natural-language queries like \u201cpull logs from 5 minutes prior\u201d or \u201cverify if this impacts user X\u201d and executes them against the connected observability stack automatically.<\/p>\n<p>This design matters for lean teams. <a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">Struct integrates with Slack, GitHub, and observability platforms for quick deployment in minutes<\/a>. The workflow engineers already use for incident communication becomes the single pane of glass for investigation, verification, and handoff.<\/p>\n<h2>Step 6: Production Safeguards and Governance<\/h2>\n<p>Struct enforces VPC-aware access, ephemeral log processing, SOC 2 Type II and HIPAA compliance, and explicit data-quality caveats so teams know when poor logging will limit accuracy. Full compliance documentation is available at trust.struct.ai.<\/p>\n<p>Data quality creates a real constraint. If a system lacks basic logging, trace IDs, or alerting triggers, automated investigation accuracy degrades. Core governance controls for self-healing ITOps include intent-based policies, blast radius limits, tiered approval paths, audit trails, and rollback capabilities to ensure automated remediation remains safe and accountable. Struct surfaces explicit caveats when data quality is insufficient instead of producing a confident but unreliable root-cause hypothesis.<\/p>\n<p>The platform\u2019s Deploy Guard feature adds instrumentation review at the pull request stage and post-deploy health checks. Teams improve alerting quality before incidents occur instead of reacting only after alerts fire.<\/p>\n<p>The six-step workflow above covers the technical mechanics. The questions below address the practical concerns teams raise when evaluating automated troubleshooting, including setup time, compliance, team readiness, and operational fit.<\/p>\n<h2>Frequently Asked Questions<\/h2>\n<h3>How long does Struct take to set up?<\/h3>\n<p>Setup takes 5 to 10 minutes. You authenticate three connection types: your issue source (Slack or PagerDuty), your code repository (GitHub), and your observability context (Datadog, Sentry, AWS CloudWatch, GCP Logs, or similar). Once connected, auto-investigations activate immediately. No professional services engagement or multi-week onboarding is required. Every Struct plan includes a 30-day risk-free pilot.<\/p>\n<h3>Is Struct SOC 2 Type II and HIPAA compliant?<\/h3>\n<p>Yes. Struct is fully SOC 2 Type II and HIPAA compliant, with documentation available at <a href=\"https:\/\/trust.struct.ai\" target=\"_blank\" rel=\"noindex nofollow\">trust.struct.ai<\/a>. Logs are accessed and processed ephemerally. Struct requires integration access to your observability tools (Datadog, AWS, GCP, and others) to function. Organizations with strict zero-egress VPC requirements that prohibit any log data leaving an internal system should evaluate the Enterprise plan\u2019s sidecar option or contact the team to discuss fit.<\/p>\n<h3>How does Struct help junior engineers own on-call?<\/h3>\n<p>Struct performs the first-pass investigation automatically, so by the time any engineer opens their laptop, the blast radius, root-cause hypothesis, and suggested fixes are already assembled in Slack. Junior engineers no longer need the tribal knowledge of a senior SRE to begin triaging an alert. Struct encodes your team\u2019s existing on-call runbooks directly into its investigation logic, so every alert response starts from the same high-quality baseline regardless of who is on call. As the Arcana case study showed, teams can scale investigation coverage significantly while maintaining investigation helpfulness above 85%.<\/p>\n<h3>What happens if our logging and telemetry are poor?<\/h3>\n<p>Struct relies on the observability data you provide. If your system lacks trace IDs, structured logs, or alerting triggers, the AI cannot deduce system state from code analysis alone. Struct surfaces explicit caveats when data quality limits investigation accuracy instead of generating a confident but unreliable output. The ideal starting point is a team already using Sentry for exceptions, Datadog or cloud logs for infrastructure, and Slack for alert routing. Struct\u2019s Deploy Guard feature helps teams improve instrumentation quality at the pull request stage before gaps create blind spots during incidents.<\/p>\n<h3>Can Struct follow our specific on-call runbooks?<\/h3>\n<p>Yes. You can input custom instructions, correlation ID formats, and your internal on-call runbook directly into Struct. The composable widget system lets you guarantee that specific visual data, such as particular dashboards, log queries, or service maps, is always pulled for defined alert types. The AI follows your exact operational procedures when an alert fires, producing outputs that match how your senior engineers would investigate the same issue.<\/p>\n<h2>Stop Burning Senior Engineers on 3 AM Log Hunting<\/h2>\n<p><a href=\"https:\/\/www.producthunt.com\/products\/struct-2\" target=\"_blank\">Large-scale customers report an 80% reduction in triage time after deploying Struct.<\/a> The time savings described across the six-step workflow compound, so investigations that once consumed 30 to 45 minutes now complete in under 5. As the Arcana case study showed, teams reclaim dozens of engineer-hours per month while maintaining investigation quality above 85%.<\/p>\n<p>Struct delivers automated troubleshooting with AI across the full six-step workflow. It handles anomaly detection, root-cause analysis, automated remediation, incident resolution verification, Slack-native collaboration, and production safeguards in under 10 minutes on top of the Datadog, Sentry, and GitHub stack your team already uses.<\/p>\n<p><a href=\"https:\/\/cal.com\/deepanm\/struct-demo\" target=\"_blank\"><strong>See the full Struct workflow in a live environment\u2014schedule a walkthrough<\/strong><\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Struct detects, diagnoses, and resolves incidents in under 5 minutes. Automate your on-call runbook and cut triage time 80%. Start your free pilot.<\/p>\n","protected":false},"author":118,"featured_media":830,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"inline_featured_image":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-831","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/posts\/831","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/comments?post=831"}],"version-history":[{"count":0,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/posts\/831\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/media\/830"}],"wp:attachment":[{"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/media?parent=831"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/categories?post=831"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/struct.ai\/articles\/wp-json\/wp\/v2\/tags?post=831"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}