Modh
AboutServicesWorkPlaybookResourcesBook a call
Book a call
Agent Skills/Workflow / Process/Bug Cleanup Triage
Workflow / Processintermediate10 min

workflow Bug Cleanup Triage

Framework for backlog bug cleanup sessions. Triage is three sequential activities — Linear hygiene, root-cause investigation, and code fixing — that must be executed in order as a hard phase gate. Includes git log pre-flight, Sentry module-tag verification, and umbrella breakdown auto-detection. Use when planning to clean up N bugs from a backlog, before dispatching research agents, or when a ticket has been In Progress for weeks without shipping.

Install this skill

git submodule add https://github.com/modh-labs/playbook.git .playbook

One submodule installs all 17 skills. Reference .playbook/skills/bug-cleanup-triage/SKILL.md from your AGENTS.md.


Bug Cleanup Triage

The Core Principle

Bug cleanup is three distinct activities that get conflated and lose efficiency when mixed:

  1. Linear hygiene — closing tickets whose fixes already shipped. Cost: seconds per ticket.
  2. Root-cause investigation — understanding why an open bug happens. Cost: 5–30 min per ticket.
  3. Code fixing — writing the fix + tests. Cost: 0.5–2 days per ticket.

These must happen in order, as a hard phase gate. Mixing them up front is the #1 waste mode in backlog cleanup sessions.

The Diagnostic Question

Before dispatching any research agent, reading any code, or writing any fix, ask:

"Have I checked git log --grep=<TICKET-ID> AND count() where module:<X> in Sentry for this ticket?"

If no → you're about to burn tokens on work that may already be done. Run Step 0 first. If yes → proceed to investigation.

Decision Tree

You have N bug tickets to clean up
  ↓
PHASE 1: LINEAR HYGIENE (cheapest, do for ALL N first)
  ├─ For each ticket ID, run: git log --all --grep="<TICKET-ID>" -i
  │   ├─ Commit found on main?
  │   │   ├─ Yes → Read commit body (git show <sha>)
  │   │   │         ↓
  │   │   │   Does the fix add Sentry/logger instrumentation?
  │   │   │         ├─ Yes → Run: count() where module:<module-name> in last 7d
  │   │   │         │         ├─ >0 events → FIX IS LIVE → close ticket (Strategy D)
  │   │   │         │         └─ 0 events  → SHIPPED BUT SILENT → do NOT close
  │   │   │         │                         Post investigation comment
  │   │   │         │                         Flag likely config gap (missing env var,
  │   │   │         │                         feature flag, upstream webhook not firing)
  │   │   │         │                         Escalate to user for config fix
  │   │   │         └─ No (pure UI fix) → close ticket with commit reference
  │   │   └─ No → proceed to Phase 2 for this ticket
  │   │
  │   └─ Also check: does this ticket have `Needs Breakdown` or `Epic` label,
  │                    >60 days in current state, ≥5 distinct symptoms in description?
  │         └─ Yes → UMBRELLA → switch to breakdown mode (see Rule 3)
  │
PHASE 2: ROOT-CAUSE INVESTIGATION (only on tickets that survived Phase 1)
  ├─ Dispatch research agent OR read code directly
  ├─ Identify root cause
  └─ Scope the fix
  ↓
PHASE 3: CODE FIXING (only on tickets with confirmed root cause)
  ├─ Write code + tests
  ├─ Run CI
  └─ Ship

Core Rules

Rule 1: Phase gate enforcement — hygiene MUST complete for the entire candidate set before any agent dispatch

WHY: Research agents cannot distinguish "this is the bug" from "this is the fix that resolved the bug" because the code comments read identically. If the fix has already shipped, the agent will describe the fix code as if it were the original bug — confidently wrong. Only git log + Sentry can distinguish.

CORRECT:

# Phase 1: hygiene pass for ALL 14 candidate tickets
for id in TICKET-901 TICKET-910 TICKET-916 TICKET-930 ...; do
  echo "=== $id ==="
  git log --oneline --all --grep="$id" -i
done

# Read any commits found. Close already-fixed tickets.
# THEN dispatch research agents only on tickets that remain.

WRONG:

# Dispatching 5 parallel research agents immediately
Task(subagent_type=Explore, prompt="Investigate TICKET-916 overbooking bug...")
Task(subagent_type=Explore, prompt="Investigate TICKET-910 status resolution...")
# ...agents return "root cause found at file:line X" — but the fix shipped yesterday
# and file:line X IS the fix, not the bug. 30 minutes wasted.

Rule 2: Shipped ≠ live — always verify with Sentry module-tag count IN THE CORRECT DATASET

WHY: A fix can be on main and still be completely silent in production. Missing env var, unset feature flag, upstream webhook not firing, silent no-op client — any of these leaves instrumented code dormant. The commit is green; the fix is not. But also: 0 events in errors means "nothing is broken," NOT "nothing is firing." Choose the dataset based on what the fix instruments.

Dataset selection matrix:

Fix emitsQuery datasetQuery shape
logger.info / logger.warn (success-path logging via a module logger)logscount of log events in the logs dataset where module equals <module-name>
captureException (error capture)errorscount of errors where stack.module:<module-name>
Both (most real fixes)logs first, then errorsRun both; >0 in logs = firing successfully, >0 in errors = firing but failing

CORRECT — verify before closing, logs dataset first:

# Fix adds `logger.info("event sent")` via createModuleLogger("meta-capi")
mcp__sentry__search_events(
  organizationSlug: "<org>",
  naturalLanguageQuery: "count of log events in the logs dataset where module equals meta-capi in the last 14 days"
)
# Returns count() = 45 → fix is live → close ticket

Result interpretation:

  • >0 in logs: fix is live, proceed to close
  • 0 in logs AND 0 in errors: fix code exists but is never reached → likely config gap (missing env var, feature flag off, upstream webhook not firing). Don't close; escalate with Sentry link and config diagnosis.
  • 0 in logs AND >0 in errors: the code path IS reached, but every invocation throws. Worse than not firing — prioritize error triage, don't close.

WRONG — trust the commit blindly:

# "I see commit abc123 on main that mentions TICKET-937. Closing."
# Reality: META_ACCESS_TOKEN is unset in Vercel prod. CAPI client is in silent
# no-op mode. Zero events have fired in 30 days. Users still can't track conversions.
# You just closed a ticket that's still actively broken.

Rule 3: Umbrella breakdown auto-detection — skip single-bug triage on umbrella tickets

WHY: A ticket labeled Needs Breakdown or Epic, stuck In Progress for >60 days, with ≥5 distinct symptoms in its description, is not one bug — it's a collection of N root causes that can't fit in a single branch. Trying to triage it as one bug will waste investigation time and produce a fix plan that can't ship. The first job is to break it into per-root-cause sub-tickets.

Auto-detect signal (ALL three must hold):

isUmbrella = (
  labels.includes("Needs Breakdown") || labels.includes("Epic")
) && (
  daysInCurrentState > 60
) && (
  countDistinctSymptomsInDescription >= 5
)

Action when umbrella detected:

  1. Do NOT dispatch research agents on the umbrella
  2. Read the full description + every "Additional report" block
  3. Propose N sub-tickets, each with: specific title, parent link, Bug/area labels, priority from symptom severity, files + root cause hypothesis + fix sketch + test plan + AC
  4. Group proposed sub-tickets into clusters by shared code path (one investigation covers multiple fixes, one PR ships multiple bugs)
  5. Post a breakdown comment on the umbrella with the full list + cluster recommendation
  6. Ask user to pick: A (keep umbrella as tracking epic + create children), B (close umbrella + create fresh with back-links), C (don't break down — almost always wrong if already stalled)

Implementation Pattern

Phase 1 — Hygiene Pass (automated, <5 min for 20 tickets)

# 1. Batch git-log check for all candidate IDs
for id in $CANDIDATE_IDS; do
  git log --oneline --all --grep="$id" -i --since="60 days ago"
done > hygiene-results.txt

# 2. For each hit, run `git show <sha>` and read the fix commit body
# 3. For each commit that adds instrumentation, run Sentry module-tag count
# 4. Classify each ticket:
#    - ALREADY FIXED (commit + Sentry >0) → Strategy D close
#    - SHIPPED BUT SILENT (commit + Sentry 0) → investigation comment, keep open
#    - NO COMMIT → survives to Phase 2
#    - UMBRELLA → breakdown mode

Phase 2 — Investigation (only on survivors)

Now you can dispatch research agents. Give each agent the specific ticket + its symptom + the hygiene-pass result ("no commit found, safe to investigate live code").

Research output gets filed as root-cause hypotheses into Phase 3.

Phase 3 — Fix

Cluster fixes by shared code path. One branch, one PR, one investigation read, one regression test suite covering multiple tickets. The cluster reduction is 3–5x — 14 tickets usually collapse into 4–6 PRs when grouped by file.

Anti-Patterns

Anti-patternWhy it failsFix
"Dispatch N research agents in parallel on day 1"Many agents run on already-fixed code and describe the fix as the bugPhase gate: hygiene pass first, agents second
"Commit on main + PR merged = ticket done"The commit may be silently dormant if config/env/feature flag is missingSentry module-tag count before closing
"The ticket description lists 9 bugs, I'll fix them all in one branch"2+ months later, branch is still open because scope is unboundedBreakdown mode: split into 9 children, ship in 3 clusters of related fixes
"I'll investigate the first ticket in depth, then move to the next"Context-switching cost. You'll re-read the same files for each ticket.Batch by cluster: read the file once, fix 3 bugs that live in it
"Close the ticket, the author said they fixed it"The commit message might say fix(X): ... but the fix could be partial, or for a different root cause than the user reportedRead the commit body AND verify the fix is firing AND match the symptom to the commit

Audit Checklist

Before marking any bug-cleanup session "done," verify:

  • Every ticket ran through git log --grep="<TICKET-ID>" first (Phase 1 hygiene)
  • Every "commit found" result had its commit body read via git show <sha> (not just title)
  • Every fix that added instrumentation was verified via Sentry module-tag count (not just "commit on main")
  • Tickets with 0 events but shipped code have an investigation comment flagging the likely config gap (env var, feature flag, upstream)
  • No research agents were dispatched on tickets that were already fixed
  • Umbrella tickets (Needs Breakdown + >60d + ≥5 symptoms) were routed to breakdown mode, not triaged as single bugs
  • Remaining tickets were clustered by shared code path and fixed in grouped PRs, not 1-ticket-per-branch
  • Each closed ticket has a Linear comment referencing the specific commit (or the Sentry verification link) — not just state: "Done"

Evidence: The Session That Taught This

2026-04-15 Aura session: 14 open bug tickets, planned as a 6-wave cleanup over multiple days.

What actually happened:

  • 9 of 14 were already fixed on main across three commits — ec418890a (AUR-901 billing), 880bbfa0b (7-bug batch), 1c794af14 (AUR-938 Nylas notify_participants).
  • 5 parallel research agents were dispatched before the git log check. They ran for 30 minutes on already-fixed code. Each agent confidently described the fix code as if it were the bug — the comment // AUR-916 insurance PUT reads identically whether you interpret it as the bug marker or the fix marker.
  • AUR-937 (Meta CAPI) — the dataset trap: code shipped correctly, META_ACCESS_TOKEN was set in Vercel production 5 days earlier, and the integration had been firing 45 events in the past 14 days. A first-pass Sentry query of count() where stack.module:meta-capi against the errors dataset returned 0 events. The first version of this skill interpreted that as "shipped but silent → escalate for config fix" and nearly closed the ticket with the wrong diagnosis. The user pushed back ("isn't the token already there?"), and a re-query against the logs dataset (where successful logger.info calls land) returned 45 events. Zero errors ≠ nothing firing. This is why Rule 2 now specifies the dataset selection matrix — the lesson was learned the hard way within hours of the skill first landing on main.
  • AUR-351: Needs Breakdown label since February, In Progress for 2+ months, description listed 9 distinct analytics bugs. Breakdown produced 9 children (AUR-951 through AUR-959) grouped into 3 clusters. First child shipped within a week.

Cost comparison:

  • Without this framework: ~30 minutes of wasted research + ~8 hours of investigation on already-fixed bugs before anyone noticed = 1 day of burned work
  • With this framework: 15 minutes of hygiene + Sentry checks, then straight to the 3 remaining real bugs = 2 hours total

The compounding lesson: Step 0 (git log + Sentry) costs seconds per ticket. Skipping it costs hours per ticket. The ROI is 100:1 in favor of checking first. Never skip.

Related Skills

workflow

Background Job Right Sizing

Choose the LIGHTEST durable mechanism a background/async/side-effect job actually needs, instead of reaching for a full workflow engine by reflex. A four-rung ladder — in-process fire-and-forget → durable queue → workflow engine → dedicated orchestrator — routed by four axes: durability, step count, concurrency shape, and cross-app/bus membership. Activates when adding any background job, emitting a side-effect from a request handler, choosing between Vercel Queues / Workflows / Inngest / Temporal / SQS, or auditing an existing job fleet for over-engineering ("do we still need the orchestrator for this?").

Workflow / Process
workflow

CI Pipeline Standards

Enforce CI pipeline conventions. Use when adding CI checks, modifying GitHub Actions workflows, or discussing CI vs deployment. Prevents deployment steps in CI and ensures the extensible step pattern is followed.

Workflow / Process
workflow

Control Flow Exceptions

Some "errors" are control flow, not failures. `redirect()` and `notFound()` in Next.js work by THROWING — a naive `try/catch` swallows the signal, cancels the navigation, and renders the internal digest string ("NEXT_REDIRECT;...") to the user. Any `catch` block wrapping code that might redirect MUST call `unstable_rethrow(error)` as its first line. Triggers on `try/catch` around a server action, Server Component, or route handler in Next.js App Router.

Workflow / Process
←Back to Agent Skills