Written by: Nimesh Chakravarthi, Co-founder & CTO, Struct
Key Takeaways
- Manual on-call triage wastes 30–45 minutes per incident as engineers jump between observability tools to find root cause.
- AI-powered automation like Struct delivers zero-click root cause analysis in under five minutes, cutting triage time by 80%.
- Struct connects Slack, PagerDuty, Datadog, Sentry, GitHub and more to correlate logs, metrics, traces, and code context automatically.
- Teams using Struct report faster MTTR, less alert noise, and junior engineers safely handling on-call shifts on their own.
- Use Struct to automate your on-call runbook, reclaim engineering time, and protect SLAs.
The Problem: Manual Triage Turns Every Incident Into a Time Sink
In a typical incident, on-call engineers often spend 15–20 minutes or more context-switching among tools such as Datadog, Slack, GitHub, and monitoring dashboards before they can begin productive debugging. Most of that time goes into investigation, not resolution.
Incident responders also face a constant stream of alerts during each shift, and most of those alerts require no immediate action. Typical enterprise security teams receive roughly 3,000–4,400 alerts per day, yet only a fraction warrant human intervention. As alert volume grows, a larger share of MTTR for production incidents gets consumed by diagnosis instead of fixes.
The downstream effects are predictable. Senior engineers become permanent firefighters. New hires cannot safely take on-call shifts without escalating nearly every alert. Product velocity slows or stalls. The New Relic 2026 AI Impact Report, drawn from 6.6 million platform users, found that AI users achieved 2x higher correlation rates and 27% less alert noise than non-AI accounts, and that gap widens as engineering teams scale. This is exactly the gap that purpose-built AI investigation platforms aim to close.
The Product: Struct Automates On-Call Investigation for SaaS Teams
Struct is an AI-powered automated on-call investigation platform built for Seed-to-Series-C SaaS engineering teams. When an alert fires in a configured Slack channel or PagerDuty queue, Struct immediately begins investigating, pulling logs, correlating metrics, mapping a timeline, and identifying root cause. By the time an engineer opens their laptop, Struct has already delivered a complete root cause report with suggested fixes in a dynamically generated dashboard.
Key capabilities include:
- Zero-click automated investigation: No prompting required. Struct triggers on alert fire and completes analysis in under five minutes.
- Slack-native conversational AI: Engineers tag Struct in any thread to pull more logs, test hypotheses, or verify blast radius without leaving Slack.
- Dynamic dashboards and unified timelines: A single view that merges Azure traces, Datadog metrics, Sentry exceptions, and GitHub commits for each incident.
- Custom runbooks and composable widgets: Teams encode their own operational procedures so Struct investigates the way a senior engineer would.
- 10-minute setup: Authenticate Slack, GitHub, and one observability source. The first automated investigation runs right away.
- SOC 2 Type II and HIPAA compliance: Struct is fully SOC 2 Type II and HIPAA compliant.
A Series A fintech with 40+ engineers reduced triage time by 80% after connecting Struct in under 10 minutes. The team protected strict SLAs and gave junior engineers enough context to handle on-call shifts without constant escalation.
See Struct automate your on-call runbook
Ranked Comparison: Triage Minutes Saved Across Tools
The table below compares reported triage investigation time before and after tool adoption. All figures come from published benchmarks or vendor-reported data cited inline.
| Tool | Avg. Triage Time (Before) | Avg. Triage Time (After) | Reduction |
|---|---|---|---|
| Struct | 45 min | <5 min | ~80% |
| PagerDuty AIOps | 2–3 hrs (MTTA) | 5 min (MTTA) | ~97% MTTA, MTTR 3 hrs → <30 min |
| Resolve AI | ~30 min (investigation phase) | <10 min | 72% (Coinbase benchmark) |
| Generic LLMs (Claude/ChatGPT) | 45 min | 30–40 min (reactive, manual log pasting required) | Minimal, no proactive investigation |
PagerDuty’s MTTA and MTTR figures reflect Anaplan’s reported results and do not match Struct’s per-investigation triage metric directly. They appear here for directional context. Resolve AI’s figures reflect Coinbase’s published benchmark across thousands of Kubernetes microservices. Generic LLMs provide no automated investigation, so engineers still gather and paste context manually, which keeps them as reactive tools instead of triage platforms.
Category Breakdown: Slack-First AI, Runbooks, and Onboarding Speed
| Category | Struct | PagerDuty | Datadog On-Call |
|---|---|---|---|
| Slack-native AI investigation | ✓ Zero-click, conversational | ✓ Bi-directional sync, GenAI summaries | Partial, monitoring-native, not Slack-first |
| Custom runbook automation | ✓ Composable widgets and runbook encoding | Workflow-based, requires configuration | Limited to Datadog monitors |
| Setup time | ~10 minutes | Hours to days (enterprise onboarding) | Requires existing Datadog instrumentation |
| New-engineer onboarding support | ✓ AI acts as automated senior engineer | Escalation routing only | Requires dashboard familiarity |
| SOC 2 / HIPAA compliance | ✓ Both | ✓ SOC 2, FedRAMP-Low (March 2025) | ✓ SOC 2 |
Struct leads on new-engineer enablement because it encodes tribal knowledge directly into the investigation output. Every alert arrives with a contextualized starting point. Engineers who are new to the system can handle on-call shifts without escalating every issue.
Struct Integrations: From Alert Trigger to Code Change
Struct’s zero-click investigation depends on deep, simultaneous access to every layer of the engineering stack. The platform connects across three integration categories.
Alert triggers: Slack channels, PagerDuty, Sentry, Linear, Jira, and Asana. When an alert fires in any configured source, Struct begins investigating immediately, with no human acknowledgment.
Observability and logs: Datadog, AWS CloudWatch, GCP Logs, Azure Logs and Traces, Grafana, Prometheus, Loki, Sumo Logic, and Better Stack. Struct queries these sources in parallel, correlates trace IDs, and merges events into a unified timeline. Struct identified a serious degradation in Slack’s web_mention webhook hours before Slack updated their own status page. That early detection came from cross-source correlation that no single observability tool surfaced alone.
Code context: GitHub. Struct cross-references recent commits and pull requests against the incident timeline to highlight deployment-correlated regressions.
Observability value comes from correlating metrics, logs, and traces quickly to answer “what’s broken and why” under pressure, not just having more data. Struct turns that principle into a default workflow by making cross-source correlation automatic instead of manual. Once root cause is confirmed, Struct hands off to a local CLI, an AI coding agent, or generates a pull request directly. This closes the loop from alert to resolution without context switching. To confirm that this automated flow delivers real value, teams then track a few concrete metrics over time.
Connect your stack and automate triage in 10 minutes
How to Measure MTTR, Alert Noise, and Ramp Time
Three core metrics show whether an on-call automation investment is working.
Triage time per incident: Measure the elapsed time from alert fire to confirmed root cause. Struct’s benchmark moves teams from 45 minutes to under 5 minutes. Track this weekly per engineer and per service so you can spot outliers and stubborn hotspots.
Alert noise ratio: Track the percentage of alerts that require human action versus those that are transient or false positives. Intelligent alert grouping and semantic deduplication reduce alert volume by 80–90% by consolidating related events across pods, clusters, and services. This metric matters because you can see that reduction only when you measure it. Struct’s automated filtering confirms which alerts require intervention and suppresses the rest, which makes this ratio visible and actionable from the first week.
New-engineer time-to-first-solo-on-call: Measure how many weeks pass before a new hire handles an on-call shift without escalating. Struct compresses this timeline by providing a contextualized investigation output for every alert. New engineers receive the same starting point a senior engineer would produce after 20 minutes of manual work.
The Catchpoint SRE Report 2026 found that SRE toil rose in 2025. Cutting triage time directly cuts toil percentage and returns engineering capacity to product work without adding headcount. The next step is deploying automation carefully so those gains appear in practice.
Pitfalls to Avoid and Struct Deployment Best Practices
Data quality is a prerequisite. Struct relies on the telemetry you provide. Teams without structured logging, trace IDs, or consistent alerting triggers will see lower investigation accuracy. At minimum, you need Sentry or equivalent exception tracking, a cloud log provider such as AWS CloudWatch, GCP, or Datadog, and Slack-based alerting.
Encode runbooks before going live. Encode runbooks first so Struct investigates according to your team’s real procedures. Struct’s composable widget system lets teams guarantee that specific data always appears for specific alert types. By copying existing on-call runbooks into Struct before the first investigation, you align outputs with your operational playbooks instead of generic patterns, which improves investigation accuracy from day one.
Compliance scoping. Struct is SOC 2 Type II and HIPAA compliant. For organizations with strict VPC egress rules that block any log data from leaving internal infrastructure, an on-premise deployment is required. Struct does not support that configuration at the Startup or Growth tiers.
Start with one high-noise channel. Connect Struct to the Slack channel with the highest alert volume first. Measure triage time before and after for two weeks, then expand to additional channels. This approach creates a clean before-and-after benchmark and builds team confidence in automated outputs before a full rollout.
Frequently Asked Questions
How much does on-call triage time improve with AI investigation tools?
Struct customers at scale report an 80% reduction in triage time, compressing standard 30–45 minute manual investigations to under five minutes. The improvement is strongest for recurring alert patterns, where Struct’s historical context matching surfaces root causes in seconds. New alert types take slightly longer while the system builds pattern recognition, but the 85–90% helpful investigation rate holds from early deployment.
Is Struct suitable for teams with strict data security requirements?
As noted in the product overview, Struct meets SOC 2 Type II and HIPAA requirements. Logs are accessed and processed ephemerally, and they are not stored beyond the investigation window. This posture covers the needs of most Seed-to-Series-C SaaS companies, including fintech and healthtech. Teams with zero-egress VPC requirements that prohibit any external log access are not currently a fit for Struct’s cloud-hosted tiers.
How quickly can an engineering team get Struct running?
Setup typically takes under 10 minutes. The process uses three connection types already described above: an alert source such as Slack or PagerDuty, a code repository such as GitHub, and an observability platform such as Datadog, AWS CloudWatch, or GCP Logs. Once connected, auto-investigations activate immediately, without professional services or lengthy onboarding.
Can Struct help junior engineers handle on-call shifts independently?
Struct acts as an automated senior engineer for the first pass of every alert. It delivers a contextualized investigation output that includes root cause, blast radius, a timeline, and a suggested fix before a human engages. This mirrors the support described in the fintech example and gives junior engineers the same starting point that previously required 20–30 minutes of senior engineer context-gathering.
What happens after Struct identifies the root cause?
Once root cause is confirmed, Struct supports a smooth handoff to resolution. Engineers can accept a suggested fix via a local CLI, pass context to an AI coding agent, or have Struct generate a pull request directly in GitHub. This flow closes the loop from alert detection to code resolution without forcing engineers to re-gather context in another tool.
Conclusion: Give Your Team Their Nights Back with Struct
Manual alert triage now counts as a solved problem. Engineering teams that still spend 30–45 minutes hunting logs across five tools at 3 a.m. pay a compounding cost in MTTR, burnout, and lost product velocity. Struct removes the investigation phase by delivering zero-click root cause in under five minutes, with an 80% reduction in triage time, a 10-minute setup, and SOC 2 and HIPAA compliance included.
Every minute saved on triage is a minute returned to shipping product. The benchmark is clear and the setup is fast. The remaining step is connecting your first alerting channel and seeing the impact on your next incident.