Modh
AboutServicesWorkPlaybookResourcesBook a call
Book a call
Agent Skills/Backend / Infrastructure/Structured Logging & Observability
Backend / Infrastructureintermediate8 min

activity Structured Logging & Observability

Enforce structured logging, Sentry integration, and distributed tracing patterns. Use when adding logging, error tracking, Sentry tags, custom spans, metrics, webhook observability, or debugging production issues. Prevents console.log usage and ensures trace correlation.

Install this skill

git submodule add https://github.com/modh-labs/playbook.git .playbook

One submodule installs all 17 skills. Reference .playbook/skills/observability/SKILL.md from your AGENTS.md.


Observability & Logging Skill

When This Skill Activates

  • Adding logging to any code (server actions, repositories, webhooks, API routes)
  • Working with error tracking or performance monitoring (e.g., Sentry)
  • Adding tracing, custom spans, or distributed trace propagation
  • Handling errors that need to be captured and alerted on
  • Discussing debugging, metrics, or alerting strategy

Decision Tree -- "What Do I Use?"

Error in webhook handler?     -> logger.error() + webhook failure capture
Error in server action?       -> logger.error() + domain-specific capture
Error in background job?      -> middleware auto-captures + logger.error()
Error in AI agent call?       -> AI exception capture (or instrumented wrapper)
Need to track duration?       -> startSpan() or webhook logger lifecycle
Need to filter in dashboard?  -> setTag() (low cardinality only)
Need debug data?              -> setContext() (not searchable)
Need charts/dashboards?       -> metrics.distribution() or .count()
Client-side error?            -> Error boundary component
Non-critical info log?        -> logger.info() (may not reach error tracker in prod)

Core Rules

1. ALWAYS Use a Structured Logger Factory -- Never console.log

// WRONG - Raw console
console.log("Processing order", orderId);

// CORRECT - Structured logger
import { createModuleLogger } from "@/lib/logger";
const logger = createModuleLogger("orders");

logger.info("Processing order", { order_id: orderId });
logger.error("Payment failed", { error, payment_id: paymentId });

Create module-specific loggers for each domain: webhookLogger, aiLogger, schedulerLogger, apiLogger, etc.

2. NEVER Double-Log -- logger.error() OR captureException(), Not Both Internally

Domain capture functions create error tracker issues (alerts). They do NOT log internally. The caller decides whether to also log for debugging:

// CORRECT -- log for debugging, then capture for alerting
logger.error("Order processing failed at payment gateway", {
  session_id: sessionId,
  stage: "payment_gateway",
});
void captureOrderException(error, {
  organizationId,
  stage: "payment_gateway",
});

// WRONG -- capture function already handles error tracker, don't also log inside it
// (causes 3-5x duplicate events)

Why this matters:

  • logger.error() -> sends to structured logs + console (deployment logs)
  • captureException() -> sends to error tracker issues (alerts, fingerprinting, grouping)
  • If both are inside the capture function AND the caller also logs, you get duplicates

3. Use Built-in OpenTelemetry -- Never Add Conflicting OTEL Packages

If your error tracking SDK has native OTEL support, use it. Adding separate OTEL packages (e.g., @vercel/otel) causes span conflicts and duplicate traces.

Only add @opentelemetry/api (lightweight API package) for trace.getActiveSpan().

4. Do Not Re-Add Console Interception After Removal

If you have removed console log interception from your error tracker config (because your logger sends directly), do NOT re-add it. It intercepts console.warn/error and re-sends them, causing duplicates with what the logger already sent.

5. Tags (Searchable) vs Context (Debug) vs Metrics (Charts)

import * as Sentry from "@sentry/nextjs"; // or your error tracker SDK

// Tags: Categorical, filterable in dashboard (low cardinality)
Sentry.setTag("webhook.source", "stripe");
Sentry.setTag("ai.verdict", "approved");

// Context: Structured debugging data (NOT searchable)
Sentry.setContext("analysis_result", {
  overallScore: 65,
  verdict: "approved",
  durationMs: 4500,
});

// Metrics: Numeric data for charts and dashboards
Sentry.metrics.distribution("ai.analysis.duration_ms", 4500, {
  unit: "millisecond",
});
Sentry.metrics.count("webhook.processed", 1, {
  attributes: { provider: "stripe" },
});
ConceptSearchable?CardinalityUse For
TagsYesLow (enum-like)Filtering issues/traces in dashboard
ContextNoAnyDebug data attached to an event
MetricsCharts onlyN/ADistributions, counts, dashboards

6. PII Auto-Redaction

Configure your logger to automatically redact sensitive fields: password, token, secret, apikey, api_key, accesstoken, access_token, refreshtoken, refresh_token, bearer, authorization, creditcard, credit_card, cardnumber, card_number, cvv, ssn, social_security, name.

Never interpolate PII into log message strings. Always pass PII as structured attributes where the redaction layer can catch it:

// CORRECT -- structured attribute (auto-redacted)
logger.info("Order created", { guest_email: email });

// WRONG -- PII in unstructured string (bypasses redaction)
logger.info(`Order created for ${email}`);

7. Log Levels Per Environment

EnvironmentMinimum LevelNotes
ProductioninfoOnly debug is filtered out
Preview/StagingdebugFull verbosity
DevelopmentdebugFull verbosity

Use logger.debug() for verbose tracing you only need locally. All info, warn, and error logs appear in production.

8. Domain-Specific Error Capture

For alertable failures (error tracker issues, not just logs), use factory-generated typed capture functions:

// Factory pattern for domain capture functions
function createDomainCapture(domain: string) {
  return (error: unknown, context: DomainContext) => {
    const normalized = normalizeError(error);
    Sentry.captureException(normalized, {
      tags: {
        [`${domain}.critical`]: "true",
        [`${domain}.stage`]: context.stage,
        [`${domain}.organization_id`]: context.organizationId,
      },
      fingerprint: [`${domain}-failure`, context.stage, context.organizationId],
    });
  };
}

// Generated capture functions per domain
const captureOrderException = createDomainCapture("order");
const capturePaymentException = createDomainCapture("payment");
const captureWebhookException = createDomainCapture("webhook");
const captureAIException = createDomainCapture("ai");

These functions: (1) normalize errors, (2) auto-inject request context tags, (3) set domain-prefixed tags, (4) apply custom fingerprinting. They do NOT log internally.

See references/sentry-patterns.md for full factory pattern, tag constants, and interfaces.

9. Centralized Tag & Metric Constants

All tag keys and metric names should be centralized constants -- never hardcode tag strings:

import { ORDER_TAGS } from "@/lib/sentry/tags";
Sentry.setTag(ORDER_TAGS.STAGE, "payment_gateway");

import { METRICS } from "@/lib/sentry/metrics";
Sentry.metrics.count(METRICS.WEBHOOK.PROCESSED, 1, { ... });

10. Span Naming Conventions

PatternExampleUse For
db.*db.orders.listDatabase operations
ai.*ai.sentiment-analysisAI/ML operations
http.*http.external-apiExternal HTTP calls
function.*function.process-paymentBusiness logic

Error Handling Patterns

Server Action / Repository Error

try {
  await processPayment(paymentId);
} catch (error) {
  logger.error("Payment processing failed", { payment_id: paymentId, error });
  void capturePaymentException(error, {
    organizationId,
    stage: "payment_storage",
    paymentIntentId,
  });
  return { success: false, error: "Payment failed" };
}

Webhook Error (Use Webhook Logger)

import { createWebhookLogger } from "@/lib/webhooks/webhook-logger";

const wLogger = createWebhookLogger({
  provider: "stripe",
  eventType,
  eventId,
});
wLogger.start();
try {
  // ... handle webhook
  wLogger.success();
} catch (error) {
  wLogger.failure(error);
  // webhook logger's failure() already calls error tracker capture
}

See references/webhook-logger.md for the full webhook logger lifecycle, handler template, and withWebhookTracing wrapper.

logError() Utility

import { logError } from "@/lib/logger";
logError("Payment processing failed", error, { payment_id: paymentId });

Custom Span Pattern

import * as Sentry from "@sentry/nextjs";

const result = await Sentry.startSpan(
  { name: "function.process-media", op: "function" },
  async (span) => {
    const result = await processMedia(input);
    span.setAttributes({ "custom.output_size": result.length });
    return result;
  }
);

Wide Events

Prefer fewer, attribute-rich logs over many thin logs:

// Wide event -- one complete, searchable log
logger.info("Order completed", {
  order_id: id,
  stage: "complete",
  duration_ms: elapsed,
  organization_id: orgId,
  outcome: "confirmed",
});

// Use logger.debug() for intermediate steps (filtered in production)

Webhook Handlers

Webhook handlers require specialized observability. Use a dedicated webhook logger factory that provides:

  1. Lifecycle methods: start() -> success() / failure()
  2. Automatic duration tracking with configurable slow-operation thresholds
  3. Early Sentry tag application for dashboard filtering
  4. Idempotency awareness -- webhooks are delivered multiple times
const log = createWebhookLogger({
  provider: "stripe",
  handler: "handlePaymentSucceeded",
  eventType: "payment_intent.succeeded",
  organizationId,
  durationThresholdMs: 5000, // warn if slower
});

log.start();  // starts timer, applies tags
try {
  // 1. Idempotency check
  const existing = await findExisting(payload.id);
  if (existing) {
    log.info({}, "Already processed (idempotent skip)");
    log.success({ reason: "already_processed" });
    return { success: true };
  }

  // 2. Business logic
  const result = await processEvent(payload);

  // 3. Update tags with created entity IDs
  log.setEntityId("order_id", result.id);

  log.success({ entity_id: result.id });
} catch (error) {
  log.failure(error);  // logs error + captures to error tracker
  throw error;
}

Full webhook logger reference, handler template, and withWebhookTracing wrapper: references/webhook-logger.md


AI Instrumentation

Use instrumented wrappers for AI agent calls with automatic spans and token tracking:

import { instrumentedAgentGenerate } from "@/lib/ai/instrumentation";

const result = await instrumentedAgentGenerate(
  agent, messages, options,
  {
    operation: "sentiment-analysis",
    modelName: "claude-sonnet-4-5-20250929",
    organizationId,
  },
);
// Auto-emits: ai.tokens.prompt, ai.tokens.completion, ai.duration metrics

For non-agent AI calls, use trackAIOperation() for lighter-weight metrics tracking.

See references/sentry-patterns.md for full AI instrumentation patterns.


Error Boundaries

Every route MUST have an error boundary component:

"use client";
import { PageErrorBoundary } from "@/components/PageErrorBoundary";

export default function Error({ error, reset }: { error: Error; reset: () => void }) {
  return <PageErrorBoundary error={error} reset={reset} title="Page" loggerModule="route-name" />;
}

Performance Thresholds

MetricWarningError
DB Query>1s>5s
Server Action>3s>10s
API Route>2s>8s
AI Agent Call>10s>30s

Anti-Patterns

Anti-PatternWhy It's WrongCorrect Alternative
console.log / console.errorNo structure, no redaction, no correlationlogger.info() / logger.error()
Direct Sentry.captureExceptionNo domain tags, no fingerprintingDomain capture*Exception() functions
Conflicting OTEL packagesSpan conflicts, duplicate tracesUse SDK built-in OTEL
Console interception integrationDuplicates with direct logger sendsRemove it
logger.error() inside capture functionCaller handles logging separatelyOnly captureException inside
High-cardinality tags (order_id, user_id)Bloats tag index, not filterableUse context instead
logger.info() for production alertsInfo is not alertableUse warn or captureException
Missing trace propagation in AI callsTraces won't connect across servicesAlways pass tracing options
PII interpolated in log stringsBypasses auto-redactionPass as structured attributes

Debugging Workflow

1. Search issues by keyword or tag
2. Get issue details -> check tags, context, stacktrace
3. Get trace details -> view span waterfall
4. Search structured logs by trace_id for full context
5. For background jobs -> copy run_id -> check job dashboard
6. Fix -> deploy -> verify issue resolves

Detailed References

  • Domain exception factory, tag constants, tracing, AI instrumentation: references/sentry-patterns.md
  • Webhook logger lifecycle, handler template, withWebhookTracing: references/webhook-logger.md

Related Skills

shield-question

Prove Your Telemetry

Instrumentation fails silently, and its failure looks exactly like good news. An absent field reads as "nothing to report", a quiet dashboard reads as "nothing is wrong", a low issue count reads as "we fixed it". Before trusting any of those, read a real emitted event and confirm it carries what you think. Use when adding or changing logging/error capture, when an error report has no useful detail, when a metric or dashboard is suspiciously quiet, when an error "stopped happening" on its own, or when a shipped fix cannot be confirmed.

Universal
shield-check

Write Criticality Classification

Classify database writes by durability requirement (tracking vs critical) to eliminate false-positive data-loss alarms. Use when adding error handling for DB writes, Sentry alerting for failed operations, or designing retry logic. Prevents noisy alerts from fire-and-forget writes while ensuring real data loss is caught and retried.

Backend / Infrastructure
copy-x

Error Report Dedup Marker

Use when two error handlers sit on the same throw path and each can report to your error tracker (e.g. an inner span/critical-path/retry wrapper that captures-and-rethrows, wrapped by an outer action wrapper that also captures). Tag the Error with a non-enumerable Symbol.for marker so the second handler skips a duplicate event. Triggers on: nested try/catch that both captureException, duplicate issues for one failure, a capture-and-rethrow wrapper, "why are we getting two Sentry events per error".

Backend / Infrastructure

Related Playbook Chapters

  • →05 Observability/structured Logging
  • →05 Observability/error Tracking
  • →05 Observability/webhook Observability
←Back to Agent Skills