Back/Engineering/Devin
AdvancedEngineering

How to Set Up Devin AI for Automated Incident Response and Triage

Connect incident alerts to a read only Devin investigation so the on call engineer arrives to a timestamped evidence packet, likely code path, and explicit unknowns instead of a cold page.

How to Set Up Devin AI for Automated Incident Response and Triage

Scott explains that Cognition pages Devin when a crash occurs, so an agent begins tracing the error before the on call engineer reaches a computer and can hand over an initial report with a suspected change and the error path.

Before you start

What you need

  • A PagerDuty or equivalent alert with service, environment, severity, and timestamps
  • Read only access to the relevant code, deploy history, logs, traces, and runbooks
  • A redaction and retention policy for logs, customer data, and credentials
  • A report template with evidence links, hypotheses, confidence, and unknowns
  • Named on call ownership and explicit authority for any production change

What you’ll make

An incident channel update that summarizes impact, correlates telemetry and recent changes, identifies the leading hypotheses with evidence, and hands the on call engineer concrete next queries without silently modifying production.

Tools used

  • Devin

    AI software engineer by Cognition Labs

Step by step

The workflow

Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.

4 steps

Step01

Integrate Devin with Your Alert System

Route a deduplicated incident event to Devin with service, environment, severity, start time, alert evidence, runbook, and on call contact. Use a dedicated read only identity and keep production mutation outside the default integration.

Step02

Automated Incident Investigation

Have Devin correlate the alert with logs, traces, deploys, feature flags, recent code changes, and known incidents. Require time bounded queries, evidence links, and alternative hypotheses instead of a single confident story.

Example prompt
Investigate incident [ID] for [service and environment] from [time window] using read only access. Summarize observed impact, timeline, relevant logs and traces, recent deploys or flag changes, and up to three root cause hypotheses with evidence for and against each. Cite exact links or query IDs, redact sensitive data, and list the next diagnostic action. Do not change production.
Step03

Receive Devin's Initial Incident Report

Post a compact report to the incident channel with impact, current status, timeline, evidence, leading hypotheses, confidence, unknowns, and recommended next queries. Notify the assigned human rather than presenting the report as a resolution.

Step04

Collaborate on Resolution

Let the on call engineer validate the evidence, choose the next diagnostic or mitigation step, and keep Devin working on bounded queries or code history. Record decisions, commands, results, and rollback signals in the incident timeline.

What good looks like

  • The report distinguishes observed evidence from hypothesis and links to the exact logs, traces, deploys, and code paths used.
  • The agent can investigate without viewing or reproducing unnecessary customer data or credentials.
  • The alert reaches the assigned human and the agent does not claim resolution from a plausible cause alone.
  • Any mitigation or production change follows the incident commander's authority, runbook, observability, and rollback requirements.

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

After the steps

Runbook notes

How to recover when the loop fails and where human judgment helps.

Recover

If it goes sideways

A broad alert starts expensive investigations for duplicate or low value events
Deduplicate by incident and service, tune severity routing, and set investigation budgets and cancellation behavior.
The agent treats the nearest recent deploy as the root cause
Label it a hypothesis, look for counterevidence, compare baseline telemetry, and confirm with a focused test or rollback signal.
Logs or traces expose credentials or personal data to the agent or incident channel
Use scoped read access, redact sensitive fields at ingestion, and keep raw evidence in approved systems.
The agent attempts a production fix beyond the current incident authority
Default the integration to investigation and reporting; require the incident specific runbook and authorized operator for any mutation.

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready