Claude Skill

log-ops

Log analysis and JSONL processing - structured extraction, cross-log correlation, timeline reconstruction, pattern search

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_log-ops-3dfaf0b.zip · 24 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/log-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Log Operations

Practical patterns for analyzing log files -- especially JSONL format used in agent conversation logs, benchmark outputs, and structured application logs.

Log Format Decision Tree

Unknown Log File
│
├─ Is it one JSON object per line?
│  ├─ Yes ──────────────────────── JSONL
│  │  ├─ Small file (<100MB)
│  │  │  └─ jq for extraction, jq -s for aggregation
│  │  ├─ Large file (100MB-1GB)
│  │  │  └─ rg prefilter then pipe to jq
│  │  └─ Huge file (>1GB)
│  │     └─ split + parallel jq, or jq --stream
│  │
│  └─ No
│     ├─ Is it one large JSON object/array?
│     │  └─ Yes ──────────────── Single JSON
│     │     └─ jq --stream for SAX-style, or jq directly if fits in memory
│     │
│     ├─ Does it have key=value pairs?
│     │  └─ Yes ──────────────── Structured (logfmt / key-value)
│     │     └─ rg for search, awk/sd for extraction, angle-grinder for aggregation
│     │
│     ├─ Does it follow syslog format? (timestamp hostname service[pid]: message)
│     │  └─ Yes ──────────────── Syslog
│     │     └─ rg for search, awk for column extraction, lnav for interactive
│     │
│     ├─ Is it space/tab delimited with consistent columns?
│     │  └─ Yes ──────────────── Column-based (access logs, CSV)
│     │     └─ awk for extraction, mlr for CSV, rg for pattern search
│     │
│     └─ Mixed or unstructured
│        └─ Plain text ─────────── Freeform
│           └─ rg for search, rg -A/-B for context, lnav for exploration

Prerequisites

Required (must be installed):

  • rg (ripgrep) - text search, prefiltering. Install: cargo install ripgrep / choco install ripgrep
  • jq - JSON/JSONL extraction and transformation. Install: brew install jq / choco install jq

Optional (enhanced capabilities, gracefully degraded without):

  • lnav - interactive log exploration with SQL queries. Install: brew install lnav / WSL: apt install lnav
  • agrind (angle-grinder) - pipeline aggregation syntax. Install: cargo install ag
  • mlr (Miller) - CSV/TSV log analysis. Install: brew install miller / choco install miller
  • GNU parallel - parallel processing of split files. Install: brew install parallel

All patterns in this skill work with just rg + jq. Optional tools add interactive exploration (lnav), pipeline aggregation (agrind), and tabular analysis (mlr).

Tool Selection Matrix

Tool Best For Speed Required?
rg (ripgrep) Raw pattern matching in any format Fastest Yes
jq JSONL structured extraction and transformation Fast Yes
jq -s JSONL aggregation (slurp all lines into array) Medium (loads all into memory) Yes (part of jq)
lnav Interactive exploration, SQL over logs Interactive Optional
agrind (angle-grinder) Pipeline aggregation and counting Fast Optional
awk Column-based log formats, field extraction Fast Pre-installed
mlr (Miller) CSV/TSV log analysis, statistics Fast Optional
fd + rg Searching across many log directories Fast Pre-installed in dev-shell
GNU parallel Splitting large files for parallel processing N/A (orchestrator) Optional

When to Use What

Need to...
│
├─ Find lines matching a pattern
│  └─ rg (always fastest for text search)
│
├─ Extract specific fields from JSONL
│  └─ jq -r '[.field1, .field2] | @tsv'
│
├─ Count/aggregate over JSONL
│  └─ jq -sc 'group_by(.field) | map({key: .[0].field, n: length})'
│
├─ Search JSONL by value then format results
│  └─ rg '"error"' file.jsonl | jq -r '.message'  (two-stage)
│
├─ Explore interactively with filtering/SQL
│  └─ lnav file.log
│
├─ Aggregate with pipeline syntax
│  └─ agrind '* | parse "* * *" as ts, level, msg | count by level'
│
├─ Extract columns from space-delimited logs
│  └─ awk '{print $1, $4, $7}' access.log
│
└─ Process CSV/TSV logs with headers
   └─ mlr --csv filter '$status >= 400' then stats1 -a count -f status

JSONL Quick Reference

The most common format for structured logs. One JSON object per line, no trailing commas, no wrapping array.

Stream Filtering (line by line, constant memory)

# Filter by field value
jq -c 'select(.level == "error")' app.jsonl

# Filter by nested field
jq -c 'select(.request.method == "POST")' app.jsonl

# Filter by multiple conditions
jq -c 'select(.level == "error" and .status >= 500)' app.jsonl

# Filter by array contains
jq -c 'select(.tags | index("critical"))' app.jsonl

# Filter by field existence
jq -c 'select(.stack_trace != null)' app.jsonl

# Negate a filter
jq -c 'select(.level != "debug")' app.jsonl

Field Extraction

# Extract single field
jq -r '.message' app.jsonl

# Extract multiple fields as TSV
jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl

# Extract with default for missing fields
jq -r '.error_code // "none"' app.jsonl

# Extract nested field safely
jq -r '.response.headers["content-type"] // "unknown"' app.jsonl

Aggregation (requires slurp: loads entire file)

# Count by field value
jq -sc 'group_by(.level) | map({level: .[0].level, count: length})' app.jsonl

# Top-N most common values
jq -sc '[.[].error_type] | group_by(.) | map({type: .[0], count: length}) | sort_by(-.count) | .[:10]' app.jsonl

# Sum a numeric field
jq -sc 'map(.duration_ms) | add' app.jsonl

# Average
jq -sc 'map(.duration_ms) | add / length' app.jsonl

# Min and max
jq -sc 'map(.duration_ms) | {min: min, max: max}' app.jsonl

Nested Extraction (agent logs, complex structures)

# Extract tool calls from conversation logs
jq -c '.content[]? | select(.type == "tool_use") | .name' conversation.jsonl

# De-escape nested JSON strings
jq -c '.content | fromjson' app.jsonl

# Flatten nested arrays
jq -c '[.events[]? | .action]' app.jsonl

# Extract from arrays of objects
jq -c '.results[]? | select(.passed == false) | {test: .name, error: .message}' results.jsonl

Two-Stage Pipeline (rg for speed, jq for structure)

# Fast prefilter then structured extraction
rg '"error"' app.jsonl | jq -r '[.timestamp, .message] | @tsv'

# Search for specific value then aggregate
rg '"timeout"' app.jsonl | jq -sc 'length'

# Pattern match then extract
rg '"user_id":"u-123"' app.jsonl | jq -c '{ts: .timestamp, action: .action}'

Time-Range Filtering

# Filter by timestamp range (ISO 8601 string comparison works)
jq -c 'select(.timestamp > "2026-03-08T10:00" and .timestamp < "2026-03-08T11:00")' app.jsonl

# Events in the last N minutes (using epoch seconds)
jq -c --arg cutoff "$(date -d '30 minutes ago' +%s)" 'select((.timestamp | sub("\\.[0-9]+Z$"; "Z") | fromdate) > ($cutoff | tonumber))' app.jsonl

# Extract hour for histogram
jq -r '.timestamp | split("T")[1] | split(":")[0]' app.jsonl | sort | uniq -c

Cross-File Join

# Extract IDs from one file, search in another
jq -r '.request_id' errors.jsonl | while read id; do
  rg "\"$id\"" responses.jsonl | jq -c '{id: .request_id, status: .status}'
done

# Faster: build lookup, then join
jq -r '.request_id' errors.jsonl | sort -u > /tmp/error_ids.txt
rg -Ff /tmp/error_ids.txt responses.jsonl | jq -c '{id: .request_id, status: .status}'

# Join two JSONL files by key using jq --slurpfile
jq --slurpfile lookup <(jq -sc 'map({(.id): .}) | add' lookup.jsonl) \
  '. + ($lookup[0][.ref_id] // {})' main.jsonl

Plain Text Log Patterns

Pattern Search with Context

# Show 5 lines before and after each match
rg -B5 -A5 "OutOfMemoryError" app.log

# Show only matching files
rg -l "FATAL" /var/log/

# Count matches per file
rg -c "ERROR" /var/log/*.log | sort -t: -k2 -rn

# Multiline patterns (stack traces)
rg -U "Exception.*\n(\s+at .*\n)+" app.log

Column Extraction with awk

# Apache/nginx access log: extract status codes
awk '{print $9}' access.log | sort | uniq -c | sort -rn

# Extract specific time range from syslog
awk '$0 >= "Mar  8 10:00" && $0 <= "Mar  8 11:00"' syslog

# Calculate average response time (column 11)
awk '{sum += $11; n++} END {print sum/n}' access.log

# Filter by status code and show URL + response time
awk '$9 >= 500 {print $7, $11"ms"}' access.log

Live Monitoring

# Follow with filtering
tail -f app.log | rg --line-buffered "ERROR"

# Follow JSONL and extract fields
tail -f app.jsonl | jq --unbuffered -r '[.timestamp, .level, .message] | @tsv'

# Follow multiple files
tail -f /var/log/service-*.log | rg --line-buffered "error|warn"

Timeline Reconstruction

Extracting and Sorting by Timestamp

# Merge multiple log files by timestamp
sort -t' ' -k1,2 service-a.log service-b.log > timeline.log

# JSONL: sort by timestamp field
jq -sc 'sort_by(.timestamp)[]' combined.jsonl > sorted.jsonl

# Extract timestamps and calculate gaps
jq -r '.timestamp' app.jsonl | awk '
  NR > 1 {
    cmd = "date -d \"" prev "\" +%s"; cmd | getline t1; close(cmd)
    cmd = "date -d \"" $0 "\" +%s"; cmd | getline t2; close(cmd)
    gap = t2 - t1
    if (gap > 5) print gap "s gap before " $0
  }
  { prev = $0 }
'

# Quick duration between first and last event
jq -sc '{start: .[0].timestamp, end: .[-1].timestamp}' app.jsonl

Calculating Durations Between Events

# Duration between paired events (start/end)
jq -sc '
  group_by(.request_id) |
  map(
    (map(select(.event == "start")) | .[0].timestamp) as $start |
    (map(select(.event == "end")) | .[0].timestamp) as $end |
    {id: .[0].request_id, start: $start, end: $end}
  )
' events.jsonl

# Identify the slowest phase
jq -sc '
  sort_by(.timestamp) |
  [range(1; length) | {
    from: .[.-1].event,
    to: .[.].event,
    gap: ((.[.].ts_epoch) - (.[.-1].ts_epoch))
  }] |
  sort_by(-.gap) | .[0]
' events.jsonl

Cross-Log Correlation

By Correlation ID

# Find a request across all service logs
fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {}

# Build a timeline for a single request
fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {} \; | jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .event] | @tsv'

By Timestamp Window

# Find events within 2 seconds of a known event
# First get the target timestamp
TARGET="2026-03-08T14:23:15"
jq -c --arg t "$TARGET" '
  select(
    .timestamp > ($t | sub("15$"; "13")) and
    .timestamp < ($t | sub("15$"; "17"))
  )
' other-service.jsonl

By Session/User

# Reconstruct a user session across log files
fd -e jsonl . /var/log/ -x rg "\"user-42\"" {} \; |
  jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .action] | @tsv'

Large File Strategies

Search Recent Only

# Last 10,000 lines (fast for append-only logs)
tail -n 10000 huge.log | rg "pattern"

# Last N lines of JSONL with structured extraction
tail -n 5000 huge.jsonl | jq -c 'select(.level == "error")'

Split for Parallel Processing

# Split into 100K-line chunks
split -l 100000 huge.jsonl /tmp/chunk_

# Process in parallel
fd 'chunk_' /tmp/ -x jq -c 'select(.level == "error")' {} > errors.jsonl

# With GNU parallel
split -l 100000 huge.jsonl /tmp/chunk_
ls /tmp/chunk_* | parallel 'jq -c "select(.level == \"error\")" {} >> /tmp/errors.jsonl'

Streaming for Huge Single JSON

# SAX-style processing of a huge JSON array
jq --stream 'select(.[0][0] == "results" and .[0][-1] == "status") | .[1]' huge.json

# Extract items from a huge array without loading all
jq -cn --stream 'fromstream(1 | truncate_stream(inputs))' huge-array.json

Two-Stage Always

# ALWAYS faster: rg filters text, jq parses survivors
rg '"error"' huge.jsonl | jq -r '.message'

# vs. SLOW: jq reads and parses every line
jq -r 'select(.level == "error") | .message' huge.jsonl

Search Across Directories

Multi-Directory Patterns

# Find all JSONL files with errors across trial directories
fd -e jsonl . trials/ -x rg -l '"error"' {}

# Count errors per log file across directories
fd -e jsonl . trials/ -x bash -c 'echo "$(rg -c "\"error\"" "$1" 2>/dev/null || echo 0) $1"' _ {}

# Extract and aggregate across directories
fd -e jsonl . trials/ -x jq -c 'select(.level == "error") | {file: input_filename, msg: .message}' {}

# Build summary table from multiple runs
for dir in trials/*/; do
  total=$(wc -l < "$dir/results.jsonl")
  errors=$(rg -c '"error"' "$dir/results.jsonl" 2>/dev/null || echo 0)
  echo -e "$dir\t$total\t$errors"
done | column -t -N DIR,TOTAL,ERRORS

Common Gotchas

Gotcha Why It Hurts Fix
jq -s on huge files loads everything into memory OOM crash or swap thrashing on files over ~500MB Use streaming: rg prefilter, jq --stream, or split + parallel
JSONL with embedded newlines in string values Line-by-line tools (rg, awk, head) split a single record across lines Use jq -c to re-compact, or jq -R 'fromjson?' to skip malformed lines
rg matches JSON keys, not just values rg "error" matches {"error_count": 0} which is not an error Use rg '"level":"error"' or pipe to jq 'select(.level == "error")'
Timezone mismatches in timestamp comparisons Events appear out of order or time ranges miss data Normalize to UTC before comparing: jq '.timestamp |= sub("\\+.*"; "Z")'
Unicode and escape sequences in log messages jq chokes on invalid UTF-8 or double-escaped strings Prefilter with rg -a (binary mode), or use jq -R for raw strings
Inconsistent JSON schemas across log lines jq errors on lines missing expected fields Use // operator for defaults: .field // "missing" and ? for optional: .arr[]?
Forgetting -c flag with jq on JSONL jq pretty-prints each line, output is no longer valid JSONL Always use jq -c when output feeds into another JSONL consumer
tail -f with jq buffering Output appears delayed or not at all Use jq --unbuffered or stdbuf -oL jq
Sorting JSONL by timestamp without slurp sort command does lexicographic sort on whole lines, not by field Either jq -sc 'sort_by(.timestamp)[]' or extract timestamp prefix first
Assuming log files are complete Logs may be rotated, compressed, or still being written Check for .gz rotated files: fd -e gz . /var/log/ -x zcat {} \| rg pattern
Single quotes in jq on Windows PowerShell/cmd do not handle single quotes the same as bash Use double quotes with escaped inner quotes, or write jq filter to a file

Reference Files

File Contents Lines
references/jsonl-patterns.md JSONL extraction, aggregation, transformation, comparison, and performance patterns ~700
references/analysis-workflows.md Agent conversation analysis, application log analysis, benchmark result parsing, cross-directory workflows ~600
references/tool-setup.md Installation and configuration for jq, lnav, angle-grinder, rg, awk, GNU parallel, Miller ~450

See Also

  • data-processing -- JSON/YAML/TOML processing with jq and yq
  • debug-ops -- Systematic debugging methodology, log-based debugging section
  • monitoring-ops -- Production observability, alerting, dashboards
  • file-search -- Finding files with fd, searching code with rg
Files (claude-mods)
  • assets
    • .gitkeep 0 B · in bundle
  • references
    • analysis-workflows.md 17.9 KB
      # Analysis Workflows Reference
      
      Practical end-to-end workflows for common log analysis tasks. Each workflow is self-contained with commands you can copy and adapt.
      
      ---
      
      ## Agent Conversation Log Analysis
      
      Claude Code and other AI agents produce JSONL conversation logs with nested content blocks. These workflows extract actionable information from those logs.
      
      ### Extract All Tool Calls
      
      ```bash
      # List every tool call in chronological order
      jq -c '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use") |
        {tool: .name, id: .id}
      ' conversation.jsonl
      
      # Count tool usage frequency
      jq -r '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use") | .name
      ' conversation.jsonl | sort | uniq -c | sort -rn
      
      # Extract tool calls with their inputs (summarized)
      jq -c '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use") |
        {
          tool: .name,
          input_preview: (.input | tostring | .[:100])
        }
      ' conversation.jsonl
      ```
      
      ### Identify What Code Was Written
      
      ```bash
      # Find all Write tool calls and extract file paths
      jq -r '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use" and .name == "Write") |
        .input.file_path
      ' conversation.jsonl
      
      # Find all Edit tool calls with file paths and old/new strings
      jq -c '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use" and .name == "Edit") |
        {file: .input.file_path, old: (.input.old_string | .[:60]), new: (.input.new_string | .[:60])}
      ' conversation.jsonl
      
      # Find all Bash commands that were run
      jq -r '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use" and .name == "Bash") |
        .input.command
      ' conversation.jsonl
      
      # Files created vs modified
      echo "=== Files Created (Write) ==="
      jq -r 'select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Write") | .input.file_path' conversation.jsonl | sort -u
      
      echo "=== Files Modified (Edit) ==="
      jq -r 'select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Edit") | .input.file_path' conversation.jsonl | sort -u
      ```
      
      ### Find Error Messages and Repeated Attempts
      
      ```bash
      # Find tool results that indicate errors
      jq -c '
        select(.role == "tool") |
        .content[]? | select(.type == "text") |
        select(.text | test("error|Error|ERROR|failed|Failed|FAILED|exception|Exception"))  |
        {text_preview: (.text | .[:200])}
      ' conversation.jsonl
      
      # Find retry patterns (same tool called multiple times with similar input)
      jq -r '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use") |
        "\(.name)\t\(.input | tostring | .[:80])"
      ' conversation.jsonl | sort | uniq -c | sort -rn | head -20
      
      # Count consecutive failures (same tool, error in result)
      jq -sc '
        [to_entries[] |
          select(.value.role == "tool") |
          {idx: .key, has_error: (.value.content | tostring | test("error|Error|failed|Failed"))}
        ] |
        map(select(.has_error)) | length
      ' conversation.jsonl
      ```
      
      ### Calculate Phase Timings
      
      ```bash
      # If messages have timestamps, calculate time between phases
      jq -sc '
        map(select(.timestamp != null)) |
        sort_by(.timestamp) |
        . as $msgs |
        {
          total_messages: length,
          first: .[0].timestamp,
          last: .[-1].timestamp,
          tool_calls: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use")] | length,
          reading_ops: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use" and (.name == "Read" or .name == "Glob" or .name == "Grep"))] | length,
          writing_ops: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use" and (.name == "Write" or .name == "Edit"))] | length,
          bash_ops: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Bash")] | length
        }
      ' conversation.jsonl
      
      # Phase breakdown by sequential grouping
      jq -c '
        select(.role == "assistant") |
        .content[]? | select(.type == "tool_use") |
        if .name == "Read" or .name == "Glob" or .name == "Grep" then "READING"
        elif .name == "Write" or .name == "Edit" then "WRITING"
        elif .name == "Bash" then "EXECUTING"
        else "OTHER"
        end
      ' conversation.jsonl | uniq -c
      ```
      
      ### Extract Thinking Blocks and Reasoning
      
      ```bash
      # Extract thinking/reasoning content
      jq -r '
        select(.role == "assistant") |
        .content[]? | select(.type == "thinking") |
        .thinking
      ' conversation.jsonl
      
      # Extract text responses (non-tool, non-thinking)
      jq -r '
        select(.role == "assistant") |
        .content[]? | select(.type == "text") |
        .text
      ' conversation.jsonl
      
      # Summary of assistant responses
      jq -c '
        select(.role == "assistant") |
        {
          has_thinking: (.content | any(.type == "thinking")),
          has_text: (.content | any(.type == "text")),
          tool_calls: [.content[]? | select(.type == "tool_use") | .name]
        }
      ' conversation.jsonl
      ```
      
      ### Build a Timeline of Actions
      
      ```bash
      # Full action timeline
      jq -r '
        if .role == "user" then
          "USER: " + (.content | if type == "string" then .[:100] else (.[] | select(.type == "text") | .text | .[:100]) end)
        elif .role == "assistant" then
          (.content[]? |
            if .type == "tool_use" then "TOOL: " + .name + " " + (.input | tostring | .[:80])
            elif .type == "text" then "TEXT: " + (.text | .[:100])
            else empty
            end
          )
        elif .role == "tool" then
          "RESULT: " + (.content | tostring | .[:100])
        else empty
        end
      ' conversation.jsonl
      
      # Condensed timeline (just tool calls and results)
      jq -c '
        if .role == "assistant" then
          .content[]? | select(.type == "tool_use") | {action: "call", tool: .name}
        elif .role == "tool" then
          {action: "result", success: (.content | tostring | test("error|Error|failed") | not)}
        else empty
        end
      ' conversation.jsonl
      ```
      
      ---
      
      ## Application Log Analysis
      
      ### Error Rate Over Time
      
      ```bash
      # Errors per minute
      jq -r 'select(.level == "error") | .timestamp | .[:16]' app.jsonl |
        sort | uniq -c
      
      # Errors per hour with total context
      jq -rsc '
        group_by(.timestamp | .[:13]) |
        map({
          hour: .[0].timestamp | .[:13],
          total: length,
          errors: (map(select(.level == "error")) | length)
        }) |
        map("\(.hour)\t\(.total)\t\(.errors)\t\(.errors * 100 / .total | round)%") |
        .[]
      ' app.jsonl | column -t -N HOUR,TOTAL,ERRORS,RATE
      
      # Error rate spike detection (>2x average)
      jq -sc '
        group_by(.timestamp | .[:13]) |
        map({hour: .[0].timestamp | .[:13], errors: (map(select(.level == "error")) | length)}) |
        (map(.errors) | add / length) as $avg |
        map(select(.errors > ($avg * 2))) |
        map("\(.hour): \(.errors) errors (avg: \($avg | round))")[]
      ' app.jsonl
      ```
      
      ### Slow Request Identification
      
      ```bash
      # Top 20 slowest requests
      jq -sc '
        sort_by(-.duration_ms) | .[:20] |
        .[] | [.timestamp, .method, .path, "\(.duration_ms)ms"] | @tsv
      ' app.jsonl | column -t
      
      # Slow requests by endpoint (p95)
      jq -sc '
        group_by(.path) |
        map({
          path: .[0].path,
          count: length,
          p50: (map(.duration_ms) | sort | .[length * 0.5 | floor]),
          p95: (map(.duration_ms) | sort | .[length * 0.95 | floor]),
          p99: (map(.duration_ms) | sort | .[length * 0.99 | floor])
        }) |
        sort_by(-.p95) | .[:10]
      ' app.jsonl
      
      # Requests exceeding SLA (e.g., 500ms)
      jq -c 'select(.duration_ms > 500) | {path, duration_ms, timestamp}' app.jsonl |
        jq -rsc 'group_by(.path) | map({path: .[0].path, count: length, worst: (map(.duration_ms) | max)}) | sort_by(-.count) | .[] | [.path, .count, .worst] | @tsv' |
        column -t -N ENDPOINT,SLA_VIOLATIONS,WORST_MS
      ```
      
      ### Error Correlation
      
      ```bash
      # Which errors occur together in the same time window?
      jq -rsc '
        map(select(.level == "error")) |
        group_by(.timestamp | .[:16]) |
        map(select(length > 1)) |
        map([.[].message] | unique | sort) |
        group_by(.) |
        map({errors: .[0], co_occurrences: length}) |
        sort_by(-.co_occurrences) | .[:10]
      ' app.jsonl
      
      # Errors that always precede another error
      jq -rsc '
        map(select(.level == "error")) |
        sort_by(.timestamp) |
        [range(1; length) | {before: .[. - 1].message, after: .[.].message}] |
        group_by([.before, .after]) |
        map({sequence: .[0], count: length}) |
        sort_by(-.count) | .[:10]
      ' app.jsonl
      
      # Error clusters (errors within 5 seconds of each other)
      jq -rsc '
        map(select(.level == "error")) |
        sort_by(.timestamp) |
        . as $errs |
        [range(1; length) |
          select(
            (($errs[.].ts_epoch // 0) - ($errs[. - 1].ts_epoch // 0)) < 5
          ) |
          {ts: $errs[.].timestamp, msg: $errs[.].message}
        ]
      ' app.jsonl
      ```
      
      ### User Session Reconstruction
      
      ```bash
      # Reconstruct a single user session
      jq -c 'select(.user_id == "user-42")' app.jsonl |
        jq -sc 'sort_by(.timestamp) | .[] | [.timestamp, .action, .path // .event] | @tsv' |
        column -t
      
      # Session summary for all users
      jq -sc '
        group_by(.user_id) |
        map({
          user: .[0].user_id,
          events: length,
          first_seen: (sort_by(.timestamp) | .[0].timestamp),
          last_seen: (sort_by(.timestamp) | .[-1].timestamp),
          unique_actions: ([.[].action] | unique | length),
          errors: (map(select(.level == "error")) | length)
        }) |
        sort_by(-.events)
      ' app.jsonl
      
      # User journey (sequence of page views)
      jq -r 'select(.user_id == "user-42" and .event == "page_view") | .path' app.jsonl
      ```
      
      ### Deployment Impact Analysis
      
      ```bash
      # Compare error rates before and after deployment
      DEPLOY_TIME="2026-03-08T14:30:00"
      
      echo "=== Before Deployment ==="
      jq -c --arg t "$DEPLOY_TIME" 'select(.timestamp < $t)' app.jsonl |
        jq -sc '{total: length, errors: (map(select(.level == "error")) | length)}'
      
      echo "=== After Deployment ==="
      jq -c --arg t "$DEPLOY_TIME" 'select(.timestamp >= $t)' app.jsonl |
        jq -sc '{total: length, errors: (map(select(.level == "error")) | length)}'
      
      # New error types after deployment
      BEFORE=$(jq -r --arg t "$DEPLOY_TIME" 'select(.timestamp < $t and .level == "error") | .message' app.jsonl | sort -u)
      AFTER=$(jq -r --arg t "$DEPLOY_TIME" 'select(.timestamp >= $t and .level == "error") | .message' app.jsonl | sort -u)
      comm -13 <(echo "$BEFORE") <(echo "$AFTER")
      
      # Response time comparison
      echo "=== Response Times Before ==="
      jq -sc --arg t "$DEPLOY_TIME" '
        map(select(.timestamp < $t and .duration_ms != null)) |
        {avg: (map(.duration_ms) | add / length | round), p95: (map(.duration_ms) | sort | .[length * 0.95 | floor])}
      ' app.jsonl
      
      echo "=== Response Times After ==="
      jq -sc --arg t "$DEPLOY_TIME" '
        map(select(.timestamp >= $t and .duration_ms != null)) |
        {avg: (map(.duration_ms) | add / length | round), p95: (map(.duration_ms) | sort | .[length * 0.95 | floor])}
      ' app.jsonl
      ```
      
      ---
      
      ## Benchmark and Test Result Analysis
      
      ### Parse Structured Test Results
      
      ```bash
      # CTRF JSON format (Common Test Report Format)
      jq -r '.results.tests[] | select(.status == "failed") | [.name, .message // "no message"] | @tsv' ctrf-report.json
      
      # CTRF summary
      jq '{
        total: .results.summary.tests,
        passed: .results.summary.passed,
        failed: .results.summary.failed,
        skipped: .results.summary.skipped,
        duration: "\(.results.summary.duration)ms"
      }' ctrf-report.json
      
      # JUnit XML (convert to JSON first with xq or yq)
      yq -p xml '.testsuites.testsuite.testcase[] | select(.failure != null) | ."+@name"' junit-results.xml
      
      # TAP (Test Anything Protocol) - extract failures
      rg "^not ok" test-output.tap | sd 'not ok \d+ - ' ''
      ```
      
      ### Compare Pass/Fail Rates Across Runs
      
      ```bash
      # Compare multiple CTRF reports
      for report in results/*/ctrf-report.json; do
        dir=$(dirname "$report" | xargs basename)
        passed=$(jq '.results.summary.passed' "$report")
        failed=$(jq '.results.summary.failed' "$report")
        total=$(jq '.results.summary.tests' "$report")
        echo -e "$dir\t$passed\t$failed\t$total"
      done | column -t -N RUN,PASSED,FAILED,TOTAL
      
      # Find tests that regressed (passed before, fail now)
      jq -r '.results.tests[] | select(.status == "passed") | .name' run1/ctrf-report.json | sort > /tmp/passed_before.txt
      jq -r '.results.tests[] | select(.status == "failed") | .name' run2/ctrf-report.json | sort > /tmp/failed_after.txt
      comm -12 /tmp/passed_before.txt /tmp/failed_after.txt
      
      # Flaky test detection (tests that flip between runs)
      for report in results/*/ctrf-report.json; do
        jq -r '.results.tests[] | "\(.name)\t\(.status)"' "$report"
      done | sort | awk -F'\t' '
        {status[$1] = status[$1] " " $2}
        END {
          for (test in status) {
            if (status[test] ~ /passed/ && status[test] ~ /failed/) {
              print "FLAKY:", test, status[test]
            }
          }
        }
      '
      ```
      
      ### Performance Regression Detection
      
      ```bash
      # Compare timing data between runs
      jq -sc '[.[] | {name: .name, duration: .duration}]' run1/results.jsonl > /tmp/run1_times.json
      jq -sc '[.[] | {name: .name, duration: .duration}]' run2/results.jsonl > /tmp/run2_times.json
      
      # Find tests that got significantly slower (>20% regression)
      jq -sc '
        [., input] |
        (.[0] | map({(.name): .duration}) | add) as $before |
        (.[1] | map({(.name): .duration}) | add) as $after |
        [$before | keys[] |
          select($after[.] != null) |
          {
            name: .,
            before: $before[.],
            after: $after[.],
            change_pct: (($after[.] - $before[.]) / $before[.] * 100 | round)
          } |
          select(.change_pct > 20)
        ] |
        sort_by(-.change_pct)
      ' /tmp/run1_times.json /tmp/run2_times.json
      
      # Aggregate metrics across trial directories
      for dir in trials/trial-*/; do
        trial=$(basename "$dir")
        if [ -f "$dir/metrics.jsonl" ]; then
          avg=$(jq -sc 'map(.duration) | add / length | round' "$dir/metrics.jsonl")
          p95=$(jq -sc 'map(.duration) | sort | .[length * 0.95 | floor]' "$dir/metrics.jsonl")
          echo -e "$trial\t$avg\t$p95"
        fi
      done | column -t -N TRIAL,AVG_MS,P95_MS
      ```
      
      ### Aggregate Metrics Across Trial Directories
      
      ```bash
      # Build summary from multiple benchmark runs
      fd -t d 'trial-' trials/ -x bash -c '
        trial=$(basename "$1")
        if [ -f "$1/results.jsonl" ]; then
          total=$(wc -l < "$1/results.jsonl")
          passed=$(jq -c "select(.passed == true)" "$1/results.jsonl" | wc -l)
          failed=$((total - passed))
          echo -e "$trial\t$total\t$passed\t$failed"
        fi
      ' _ {} | sort | column -t -N TRIAL,TOTAL,PASSED,FAILED
      
      # Combine all results into one file with trial label
      fd -t d 'trial-' trials/ -x bash -c '
        trial=$(basename "$1")
        jq -c --arg trial "$trial" ". + {trial: \$trial}" "$1/results.jsonl"
      ' _ {} > combined_results.jsonl
      
      # Then aggregate across all trials
      jq -sc '
        group_by(.trial) |
        map({
          trial: .[0].trial,
          total: length,
          pass_rate: ((map(select(.passed == true)) | length) / length * 100 | round),
          avg_duration: (map(.duration) | add / length | round)
        }) |
        sort_by(.trial)
      ' combined_results.jsonl
      ```
      
      ---
      
      ## Cross-Directory Analysis
      
      ### Search Pattern Across All Log Directories
      
      ```bash
      # Find which log files contain a specific error
      fd -e jsonl -e log . /var/log/services/ -x rg -l "ConnectionTimeout" {}
      
      # Count occurrences per directory
      fd -e jsonl . logs/ -x bash -c '
        count=$(rg -c "error" "$1" 2>/dev/null || echo 0)
        echo -e "$(dirname "$1" | xargs basename)\t$(basename "$1")\t$count"
      ' _ {} | sort -t$'\t' -k3 -rn | column -t -N DIR,FILE,ERRORS
      
      # Search for a pattern and show matching lines with source file
      fd -e jsonl . logs/ -x bash -c '
        rg "\"error\"" "$1" 2>/dev/null | while read line; do
          echo "$1: $line"
        done
      ' _ {}
      ```
      
      ### Build Summary Table from Multiple Log Files
      
      ```bash
      # Summary statistics per log file
      echo -e "FILE\tLINES\tERRORS\tWARNS\tFIRST_TS\tLAST_TS"
      fd -e jsonl . logs/ | while read f; do
        lines=$(wc -l < "$f")
        errors=$(rg -c '"error"' "$f" 2>/dev/null || echo 0)
        warns=$(rg -c '"warn"' "$f" 2>/dev/null || echo 0)
        first=$(head -1 "$f" | jq -r '.timestamp // "unknown"')
        last=$(tail -1 "$f" | jq -r '.timestamp // "unknown"')
        echo -e "$(basename "$f")\t$lines\t$errors\t$warns\t$first\t$last"
      done | column -t
      
      # Health check across all services
      fd -e jsonl -d 1 . /var/log/services/ -x bash -c '
        svc=$(basename "$1" .jsonl)
        last_error=$(tac "$1" | jq -r "select(.level == \"error\") | .timestamp" 2>/dev/null | head -1)
        error_count=$(rg -c "\"error\"" "$1" 2>/dev/null || echo 0)
        echo -e "$svc\t$error_count\t${last_error:-none}"
      ' _ {} | sort | column -t -N SERVICE,ERRORS,LAST_ERROR
      ```
      
      ### Identify Common Failure Patterns Across Runs
      
      ```bash
      # Extract all error messages across trial directories
      fd -e jsonl . trials/ -x jq -r 'select(.level == "error") | .message' {} |
        sort | uniq -c | sort -rn | head -20
      
      # Find which trials share the same failure
      fd -e jsonl . trials/ -x bash -c '
        trial=$(echo "$1" | rg -o "trial-[^/]+")
        jq -r "select(.level == \"error\") | .message" "$1" 2>/dev/null |
          while read msg; do echo -e "$trial\t$msg"; done
      ' _ {} |
        sort -t$'\t' -k2 |
        awk -F'\t' '
          prev != $2 { if (NR > 1 && count > 1) print count, prev_msg, trials; count=0; trials="" }
          { count++; trials = trials " " $1; prev = $2; prev_msg = $2 }
          END { if (count > 1) print count, prev_msg, trials }
        ' | sort -rn | head -10
      
      # Correlation: which errors appear together
      fd -e jsonl . trials/ -x bash -c '
        trial=$(echo "$1" | rg -o "trial-[^/]+")
        errors=$(jq -r "select(.level == \"error\") | .message" "$1" 2>/dev/null | sort -u | paste -sd "|")
        [ -n "$errors" ] && echo -e "$trial\t$errors"
      ' _ {} | sort -t$'\t' -k2 | uniq -f1 -c | sort -rn
      ```
      
      ### fd + rg + jq Composition
      
      ```bash
      # The canonical three-stage pipeline for multi-directory log analysis:
      # 1. fd: find the files
      # 2. rg: prefilter for speed
      # 3. jq: structured extraction
      
      # Example: find all timeout errors across services, extract details
      fd -e jsonl . /var/log/ |                          # find log files
        xargs rg -l '"timeout"' |                        # filter to files with timeouts
        xargs -I{} jq -c '
          select(.message | test("timeout")) |
          {file: input_filename, ts: .timestamp, svc: .service, msg: .message}
        ' {}
      
      # Example: aggregate error counts by service across all log directories
      fd -e jsonl . /var/log/ -x rg -c '"error"' {} |    # count errors per file
        awk -F: '{
          split($1, parts, "/")
          svc = parts[length(parts)-1]
          gsub(/\.jsonl$/, "", svc)
          sum[svc] += $2
        }
        END { for (s in sum) print sum[s], s }' |
        sort -rn
      
      # Example: find the most recent error across all services
      fd -e jsonl . /var/log/ -x tail -1 {} |            # last line of each file
        jq -sc '
          map(select(.level == "error")) |
          sort_by(.timestamp) |
          .[-1] |
          {service: .service, timestamp: .timestamp, message: .message}
        '
      ```
      
    • jsonl-patterns.md 17 KB
      # JSONL Patterns Reference
      
      Comprehensive patterns for working with JSONL (JSON Lines) files -- one JSON object per line, the dominant format for structured logs, agent conversation records, and streaming data.
      
      ## JSONL Basics
      
      ### Format Rules
      
      - One valid JSON object per line
      - No trailing commas between lines
      - No wrapping array or outer object
      - Each line is independently parseable
      - Newlines within string values must be escaped as `\n`
      
      ### Streaming vs Slurp
      
      ```bash
      # STREAMING (default): processes one line at a time, constant memory
      jq -c 'select(.level == "error")' app.jsonl
      
      # SLURP (-s): loads ALL lines into a single array, requires memory for entire file
      jq -sc 'group_by(.level)' app.jsonl
      
      # Rule of thumb:
      #   File < 100MB  --> slurp is fine
      #   File 100MB-1GB --> slurp with caution, prefer streaming + sort/uniq
      #   File > 1GB    --> never slurp, use streaming or split+parallel
      ```
      
      ### Key jq Flags for JSONL
      
      | Flag | Purpose | Example |
      |------|---------|---------|
      | `-c` | Compact output (one line per object) | `jq -c '.' file.jsonl` |
      | `-r` | Raw string output (no quotes) | `jq -r '.message' file.jsonl` |
      | `-s` | Slurp all lines into array | `jq -s 'length' file.jsonl` |
      | `-e` | Exit with error if output is false/null | `jq -e '.status == 200' line.json` |
      | `-R` | Read each line as raw string | `jq -R 'fromjson? // empty' messy.jsonl` |
      | `--stream` | SAX-style path/value pairs | `jq --stream '.' huge.json` |
      | `--slurpfile` | Load a file as variable | `jq --slurpfile ids ids.json 'select(.id | IN($ids[][]))' data.jsonl` |
      | `--arg` | Pass string variable | `jq --arg name "foo" 'select(.name == $name)' data.jsonl` |
      | `--argjson` | Pass JSON variable | `jq --argjson min 100 'select(.count > $min)' data.jsonl` |
      | `--unbuffered` | Flush output after each line | `tail -f app.jsonl \| jq --unbuffered -r '.message'` |
      
      ---
      
      ## Extraction Patterns
      
      ### Select by Field Value
      
      ```bash
      # Exact match
      jq -c 'select(.level == "error")' app.jsonl
      
      # Numeric comparison
      jq -c 'select(.status >= 400)' app.jsonl
      
      # String contains
      jq -c 'select(.message | test("timeout"))' app.jsonl
      
      # Regex match
      jq -c 'select(.path | test("^/api/v[0-9]+/users"))' app.jsonl
      
      # Case-insensitive match
      jq -c 'select(.message | test("error"; "i"))' app.jsonl
      
      # Null check
      jq -c 'select(.error != null)' app.jsonl
      
      # Boolean field
      jq -c 'select(.retry == true)' app.jsonl
      ```
      
      ### Select by Nested Field
      
      ```bash
      # Dot notation for nesting
      jq -c 'select(.request.method == "POST")' app.jsonl
      
      # Deep nesting
      jq -c 'select(.context.user.role == "admin")' app.jsonl
      
      # Safe navigation (no error if path missing)
      jq -c 'select(.request?.headers?["authorization"] != null)' app.jsonl
      ```
      
      ### Select by Array Contains
      
      ```bash
      # Array contains value
      jq -c 'select(.tags | index("critical"))' app.jsonl
      
      # Any element matches condition
      jq -c 'select(.events | any(.type == "error"))' app.jsonl
      
      # All elements match condition
      jq -c 'select(.checks | all(.passed == true))' app.jsonl
      
      # Array length
      jq -c 'select((.retries | length) > 3)' app.jsonl
      ```
      
      ### Extract and Flatten Nested Structures
      
      ```bash
      # Flatten one level of nesting
      jq -c '{timestamp, level, msg: .message, user: .context.user.id}' app.jsonl
      
      # Explode array into separate lines
      jq -c '.events[]' app.jsonl
      
      # Flatten array with parent context
      jq -c '. as $parent | .events[] | {request_id: $parent.request_id, event: .type, ts: .timestamp}' app.jsonl
      
      # Extract from array of objects
      jq -c '.results[] | select(.score < 0.5) | {name, score}' results.jsonl
      
      # Recursive descent (find all values for a key at any depth)
      jq -c '.. | .error_message? // empty' app.jsonl
      ```
      
      ### Handle Optional and Nullable Fields
      
      ```bash
      # Default value for missing field
      jq -r '.region // "unknown"' app.jsonl
      
      # Default for nested missing field
      jq -r '.response.body.error // .response.status_text // "no error info"' app.jsonl
      
      # Skip lines where field is missing (instead of outputting null)
      jq -r '.optional_field // empty' app.jsonl
      
      # Coalesce multiple possible fields
      jq -r '(.error_message // .err_msg // .error // "none")' app.jsonl
      
      # Type check before access
      jq -c 'if .data | type == "array" then .data | length else 0 end' app.jsonl
      ```
      
      ### Multi-Level Nesting (Agent Conversation Logs)
      
      ```bash
      # Claude Code conversation logs have deeply nested tool calls
      # Structure: {role, content: [{type: "tool_use", name, input}, ...]}
      
      # Extract all tool call names
      jq -c '.content[]? | select(.type == "tool_use") | .name' conversation.jsonl
      
      # Extract tool inputs
      jq -c '.content[]? | select(.type == "tool_use") | {tool: .name, input: .input}' conversation.jsonl
      
      # Extract text content blocks
      jq -r '.content[]? | select(.type == "text") | .text' conversation.jsonl
      
      # Extract tool results
      jq -c '.content[]? | select(.type == "tool_result") | {tool_use_id, content}' conversation.jsonl
      
      # Find tool calls that contain specific patterns in their input
      jq -c '.content[]? | select(.type == "tool_use" and (.input | tostring | test("SELECT")))' conversation.jsonl
      ```
      
      ### De-Escape Nested JSON Strings
      
      ```bash
      # When a field contains a JSON string that needs parsing
      jq -c '.payload | fromjson' app.jsonl
      
      # Safe de-escape (skip if not valid JSON)
      jq -c '.payload | fromjson? // {raw: .}' app.jsonl
      
      # Double-escaped JSON (escaped twice)
      jq -c '.data | fromjson | fromjson' app.jsonl
      
      # Extract field from de-escaped nested JSON
      jq -r '.payload | fromjson | .result.status' app.jsonl
      
      # Handle mixed escaped/unescaped
      jq -c 'if (.payload | type) == "string" then .payload | fromjson else .payload end' app.jsonl
      ```
      
      ---
      
      ## Aggregation Patterns
      
      All aggregation patterns use `-s` (slurp) which loads the entire file into memory. For large files, prefilter with `rg` first.
      
      ### Count by Field Value
      
      ```bash
      # Count per level
      jq -sc 'group_by(.level) | map({level: .[0].level, count: length})' app.jsonl
      
      # Count per status code
      jq -sc 'group_by(.status) | map({status: .[0].status, count: length}) | sort_by(-.count)' app.jsonl
      
      # Count unique values
      jq -sc '[.[].user_id] | unique | length' app.jsonl
      
      # Frequency distribution
      jq -rsc 'group_by(.level) | map("\(.[0].level)\t\(length)") | .[]' app.jsonl
      ```
      
      ### Sum, Average, Min, Max
      
      ```bash
      # Sum
      jq -sc 'map(.bytes) | add' app.jsonl
      
      # Average
      jq -sc 'map(.duration_ms) | add / length' app.jsonl
      
      # Min and max
      jq -sc 'map(.duration_ms) | {min: min, max: max, avg: (add / length)}' app.jsonl
      
      # Percentile approximation (p50, p95, p99)
      jq -sc '
        map(.duration_ms) | sort |
        length as $n |
        {
          p50: .[($n * 0.50 | floor)],
          p95: .[($n * 0.95 | floor)],
          p99: .[($n * 0.99 | floor)],
          max: .[-1]
        }
      ' app.jsonl
      
      # Sum grouped by category
      jq -sc '
        group_by(.service) |
        map({service: .[0].service, total_bytes: (map(.bytes) | add)})
      ' app.jsonl
      ```
      
      ### Group By with Aggregation
      
      ```bash
      # Group by service, show count and error rate
      jq -sc '
        group_by(.service) |
        map({
          service: .[0].service,
          total: length,
          errors: (map(select(.level == "error")) | length),
          error_rate: ((map(select(.level == "error")) | length) / length * 100 | round)
        })
      ' app.jsonl
      
      # Group by hour
      jq -sc '
        group_by(.timestamp | split("T")[1] | split(":")[0]) |
        map({
          hour: .[0].timestamp | split("T")[1] | split(":")[0],
          count: length
        })
      ' app.jsonl
      
      # Nested group by (service then level)
      jq -sc '
        group_by(.service) |
        map({
          service: .[0].service,
          by_level: (group_by(.level) | map({level: .[0].level, n: length}))
        })
      ' app.jsonl
      ```
      
      ### Top-N Queries
      
      ```bash
      # Top 10 slowest requests
      jq -sc 'sort_by(-.duration_ms) | .[:10] | .[] | {path: .path, ms: .duration_ms}' app.jsonl
      
      # Top 5 most frequent errors
      jq -sc '
        map(select(.level == "error")) |
        group_by(.message) |
        map({message: .[0].message, count: length}) |
        sort_by(-.count) | .[:5]
      ' app.jsonl
      
      # Top users by request count
      jq -sc '
        group_by(.user_id) |
        map({user: .[0].user_id, requests: length}) |
        sort_by(-.requests) | .[:10]
      ' app.jsonl
      ```
      
      ### Histogram and Distribution Analysis
      
      ```bash
      # Response time histogram (buckets: 0-100, 100-500, 500-1000, 1000+)
      jq -sc '
        map(.duration_ms) |
        {
          "0-100ms": (map(select(. < 100)) | length),
          "100-500ms": (map(select(. >= 100 and . < 500)) | length),
          "500-1000ms": (map(select(. >= 500 and . < 1000)) | length),
          "1000ms+": (map(select(. >= 1000)) | length)
        }
      ' app.jsonl
      
      # Status code distribution
      jq -rsc '
        group_by(.status) |
        map("\(.[0].status)\t\(length)") |
        sort | .[]
      ' app.jsonl
      
      # Log level distribution over time (by hour)
      jq -rsc '
        group_by(.timestamp | split("T")[1] | split(":")[0]) |
        map(
          (.[0].timestamp | split("T")[1] | split(":")[0]) as $hour |
          {
            hour: $hour,
            info: (map(select(.level == "info")) | length),
            warn: (map(select(.level == "warn")) | length),
            error: (map(select(.level == "error")) | length)
          }
        ) | .[] | [.hour, .info, .warn, .error] | @tsv
      ' app.jsonl | column -t -N HOUR,INFO,WARN,ERROR
      ```
      
      ### Running Totals and Cumulative Sums
      
      ```bash
      # Cumulative error count over time
      jq -sc '
        sort_by(.timestamp) |
        reduce .[] as $item (
          {total: 0, rows: []};
          .total += 1 |
          .rows += [{ts: $item.timestamp, cumulative: .total}]
        ) | .rows[] | [.ts, .cumulative] | @tsv
      ' <(jq -c 'select(.level == "error")' app.jsonl)
      
      # Running average of response times
      jq -sc '
        sort_by(.timestamp) |
        foreach .[] as $item (
          {n: 0, sum: 0};
          .n += 1 | .sum += $item.duration_ms;
          {ts: $item.timestamp, running_avg: (.sum / .n | round)}
        )
      ' app.jsonl
      ```
      
      ---
      
      ## Transformation Patterns
      
      ### Reshape Objects
      
      ```bash
      # Flatten nested to flat
      jq -c '{
        ts: .timestamp,
        level: .level,
        msg: .message,
        user: .context.user.id,
        method: .request.method,
        path: .request.path
      }' app.jsonl
      
      # Add computed fields
      jq -c '. + {
        date: (.timestamp | split("T")[0]),
        hour: (.timestamp | split("T")[1] | split(":")[0] | tonumber),
        is_error: (.level == "error")
      }' app.jsonl
      
      # Rename fields
      jq -c '{timestamp: .ts, message: .msg, severity: .lvl}' app.jsonl
      
      # Remove fields
      jq -c 'del(.stack_trace, .internal_debug_info)' app.jsonl
      ```
      
      ### Merge Fields from Multiple Lines
      
      ```bash
      # Combine start and end events by request_id
      jq -sc '
        group_by(.request_id) |
        map(
          (map(select(.event == "start")) | .[0]) as $start |
          (map(select(.event == "end")) | .[0]) as $end |
          {
            request_id: .[0].request_id,
            start: $start.timestamp,
            end: $end.timestamp,
            status: $end.status,
            path: $start.path
          }
        )[]
      ' events.jsonl
      
      # Merge consecutive lines (e.g., multiline log entries)
      jq -sc '
        reduce .[] as $item (
          [];
          if (. | length) == 0 then [$item]
          elif $item.continuation == true then
            (.[-1].message += "\n" + $item.message) | .
          else . + [$item]
          end
        )[]
      ' app.jsonl
      ```
      
      ### Convert Between Formats
      
      ```bash
      # JSONL to CSV
      jq -r '[.timestamp, .level, .message] | @csv' app.jsonl > app.csv
      
      # JSONL to TSV
      jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl > app.tsv
      
      # JSONL to CSV with header
      echo "timestamp,level,message" > app.csv
      jq -r '[.timestamp, .level, .message] | @csv' app.jsonl >> app.csv
      
      # CSV to JSONL (using mlr)
      mlr --c2j cat app.csv > app.jsonl
      
      # JSONL to formatted table
      jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl | column -t -s$'\t'
      
      # JSONL to markdown table
      echo "| Timestamp | Level | Message |"
      echo "|-----------|-------|---------|"
      jq -r '"| \(.timestamp) | \(.level) | \(.message) |"' app.jsonl
      ```
      
      ### Annotate Lines with Computed Fields
      
      ```bash
      # Add line number
      jq -c --argjson n 0 '. + {line_num: (input_line_number)}' app.jsonl
      
      # Add duration since previous event (requires slurp)
      jq -sc '
        sort_by(.timestamp) |
        . as $all |
        [range(length)] |
        map(
          $all[.] + (
            if . > 0 then {gap_from_prev: "computed"}
            else {gap_from_prev: null}
            end
          )
        )[]
      ' app.jsonl
      
      # Tag lines matching criteria
      jq -c '. + {
        severity_class: (
          if .level == "error" or .level == "fatal" then "critical"
          elif .level == "warn" then "warning"
          else "normal"
          end
        )
      }' app.jsonl
      
      # Enrich with filename when processing multiple files
      fd -e jsonl . logs/ -x bash -c 'jq -c --arg src "$1" ". + {source: \$src}" "$1"' _ {}
      ```
      
      ---
      
      ## Comparison Patterns
      
      ### Diff Two JSONL Files by Matching Key
      
      ```bash
      # Find entries in A but not in B (by id)
      jq -r '.id' b.jsonl | sort > /tmp/b_ids.txt
      jq -c --slurpfile bids <(jq -Rs 'split("\n") | map(select(. != ""))' /tmp/b_ids.txt) '
        select(.id | IN($bids[0][]))  | not
      ' a.jsonl
      
      # Simpler approach using comm
      jq -r '.id' a.jsonl | sort > /tmp/a_ids.txt
      jq -r '.id' b.jsonl | sort > /tmp/b_ids.txt
      comm -23 /tmp/a_ids.txt /tmp/b_ids.txt  # IDs in A but not B
      comm -13 /tmp/a_ids.txt /tmp/b_ids.txt  # IDs in B but not A
      comm -12 /tmp/a_ids.txt /tmp/b_ids.txt  # IDs in both
      
      # Find records that exist in both but have different values
      jq -sc '
        [., input] |
        (.[0] | map({(.id): .}) | add) as $a |
        (.[1] | map({(.id): .}) | add) as $b |
        ($a | keys) as $keys |
        [$keys[] | select($a[.] != $b[.])] |
        map({id: ., a: $a[.], b: $b[.]})
      ' <(jq -sc '.' a.jsonl) <(jq -sc '.' b.jsonl)
      ```
      
      ### Side-by-Side Field Comparison
      
      ```bash
      # Compare a specific field between two runs
      paste <(jq -r '[.id, .score] | @tsv' run1.jsonl | sort) \
            <(jq -r '[.id, .score] | @tsv' run2.jsonl | sort) |
        awk -F'\t' '$2 != $4 {print $1, "run1=" $2, "run2=" $4}'
      
      # Summary comparison of two log files
      echo "=== File A ===" && jq -sc '{
        lines: length,
        errors: (map(select(.level == "error")) | length),
        unique_users: ([.[].user_id] | unique | length)
      }' a.jsonl
      echo "=== File B ===" && jq -sc '{
        lines: length,
        errors: (map(select(.level == "error")) | length),
        unique_users: ([.[].user_id] | unique | length)
      }' b.jsonl
      ```
      
      ### Find New, Missing, and Changed Records
      
      ```bash
      # Comprehensive diff report
      jq -r '.id' a.jsonl | sort > /tmp/a.ids
      jq -r '.id' b.jsonl | sort > /tmp/b.ids
      
      echo "--- New in B (not in A) ---"
      comm -13 /tmp/a.ids /tmp/b.ids
      
      echo "--- Removed from A (not in B) ---"
      comm -23 /tmp/a.ids /tmp/b.ids
      
      echo "--- Changed (in both, different values) ---"
      comm -12 /tmp/a.ids /tmp/b.ids | while read id; do
        a_hash=$(rg "\"id\":\"$id\"" a.jsonl | md5sum | cut -d' ' -f1)
        b_hash=$(rg "\"id\":\"$id\"" b.jsonl | md5sum | cut -d' ' -f1)
        [ "$a_hash" != "$b_hash" ] && echo "$id"
      done
      ```
      
      ---
      
      ## Performance Patterns
      
      ### Two-Stage rg + jq Pipeline
      
      The single most important performance pattern. ripgrep is 10-100x faster than jq at scanning text.
      
      ```bash
      # BAD: jq scans every line (slow on large files)
      jq -c 'select(.level == "error" and .service == "auth")' huge.jsonl
      
      # GOOD: rg filters text first, jq only parses matching lines
      rg '"error"' huge.jsonl | rg '"auth"' | jq -c '.'
      
      # GOOD: for precise matching after rg prefilter
      rg '"error"' huge.jsonl | jq -c 'select(.level == "error" and .service == "auth")'
      
      # Benchmarks (typical 1GB JSONL file):
      #   jq alone:     45 seconds
      #   rg + jq:      3 seconds
      #   rg alone:     0.8 seconds
      ```
      
      ### GNU parallel for Splitting Large Files
      
      ```bash
      # Split a 10GB file and process in parallel
      split -l 500000 huge.jsonl /tmp/chunk_
      
      # Count errors across all chunks
      ls /tmp/chunk_* | parallel "rg -c '\"error\"' {}" | awk -F: '{sum+=$2} END {print sum}'
      
      # Extract and merge results
      ls /tmp/chunk_* | parallel "jq -c 'select(.level == \"error\")' {}" > all_errors.jsonl
      
      # Cleanup
      rm /tmp/chunk_*
      
      # One-liner with process substitution
      parallel --pipe -L 100000 'jq -c "select(.level == \"error\")"' < huge.jsonl > errors.jsonl
      ```
      
      ### jq --stream for SAX-Style Processing
      
      For files too large to fit in memory, even line-by-line (e.g., a single 5GB JSON array).
      
      ```bash
      # Count items in a huge JSON array without loading it
      jq --stream 'select(.[0] | length == 1) | .[0][0]' huge-array.json | tail -1
      
      # Extract specific field from each item in huge array
      jq -cn --stream 'fromstream(1 | truncate_stream(inputs)) | .name' huge-array.json
      
      # Filter items from huge array
      jq -cn --stream '
        fromstream(1 | truncate_stream(inputs)) |
        select(.status == "failed")
      ' huge-array.json
      ```
      
      ### Indexing Frequently-Queried Files
      
      ```bash
      # Build an index of line offsets by key value
      awk '{
        match($0, /"request_id":"([^"]+)"/, m)
        if (m[1]) print m[1], NR
      }' app.jsonl | sort > app.idx
      
      # Look up specific request by index
      LINE=$(grep "req-abc-123" app.idx | awk '{print $2}')
      sed -n "${LINE}p" app.jsonl | jq .
      
      # Build a SQLite index for repeated queries
      sqlite3 log_index.db "CREATE TABLE idx (request_id TEXT, line INTEGER)"
      awk '{
        match($0, /"request_id":"([^"]+)"/, m)
        if (m[1]) print "INSERT INTO idx VALUES (\047" m[1] "\047, " NR ");"
      }' app.jsonl | sqlite3 log_index.db
      
      # Query by index
      LINE=$(sqlite3 log_index.db "SELECT line FROM idx WHERE request_id = 'req-abc-123'")
      sed -n "${LINE}p" app.jsonl | jq .
      ```
      
      ### Memory-Efficient Aggregation Without Slurp
      
      ```bash
      # Count by level without loading entire file
      jq -r '.level' app.jsonl | sort | uniq -c | sort -rn
      
      # Top error messages without slurp
      jq -r 'select(.level == "error") | .message' app.jsonl | sort | uniq -c | sort -rn | head -20
      
      # Unique users without slurp
      jq -r '.user_id' app.jsonl | sort -u | wc -l
      
      # Sum without slurp
      jq -r '.bytes' app.jsonl | awk '{sum+=$1} END {print sum}'
      
      # These are all O(1) memory (streaming) vs O(n) memory (slurp)
      ```
      
    • tool-setup.md 17.9 KB
      # Tool Setup Reference
      
      Installation, configuration, and key commands for log analysis tools. Each tool includes install commands for all platforms, the most useful flags, and integration patterns.
      
      ---
      
      ## jq -- JSON/JSONL Processor
      
      The primary tool for structured log analysis. Processes JSONL line by line (streaming) or as a batch (slurp).
      
      ### Installation
      
      ```bash
      # macOS
      brew install jq
      
      # Ubuntu/Debian
      sudo apt install jq
      
      # Windows
      choco install jq
      # or
      winget install jqlang.jq
      
      # Verify
      jq --version
      ```
      
      ### Key Flags
      
      | Flag | Purpose | Example |
      |------|---------|---------|
      | `-c` | Compact output (one JSON per line) | `jq -c '.' file.jsonl` |
      | `-r` | Raw string output (no quotes) | `jq -r '.message' file.jsonl` |
      | `-s` | Slurp: read all lines into array | `jq -s 'length' file.jsonl` |
      | `-S` | Sort object keys | `jq -S '.' file.json` |
      | `-e` | Exit 1 if output is false/null | `jq -e '.ok' file.json` |
      | `-R` | Read lines as raw strings | `jq -R 'fromjson?' messy.jsonl` |
      | `-n` | Null input (use with inputs) | `jq -n '[inputs]' file.jsonl` |
      | `--arg` | Pass string variable | `jq --arg id "42" 'select(.id == $id)'` |
      | `--argjson` | Pass JSON variable | `jq --argjson n 10 'select(.count > $n)'` |
      | `--slurpfile` | Load file as variable | `jq --slurpfile ids ids.json 'select(.id | IN($ids[][]))'` |
      | `--stream` | SAX-style path/value output | `jq --stream '.' huge.json` |
      | `--unbuffered` | Flush after each output | `tail -f f.jsonl \| jq --unbuffered '.'` |
      | `--tab` | Use tabs for indentation | `jq --tab '.' file.json` |
      
      ### Essential Commands
      
      ```bash
      # Pretty print a single JSON object
      jq '.' file.json
      
      # Validate JSONL (report bad lines)
      jq -c '.' file.jsonl > /dev/null 2>&1 || echo "Invalid JSON detected"
      
      # Find and show invalid lines
      awk '{
        cmd = "echo " "'\'''" $0 "'\'''" " | jq . 2>/dev/null"
        if (system(cmd) != 0) print NR": "$0
      }' file.jsonl
      
      # Better: use jq -R to find invalid lines
      jq -R 'fromjson? // error' file.jsonl 2>&1 | rg "error" | head
      
      # Count lines in JSONL
      jq -sc 'length' file.jsonl
      
      # Get unique keys across all objects
      jq -sc '[.[] | keys[]] | unique' file.jsonl
      
      # Get schema (keys and types) from first line
      head -1 file.jsonl | jq '[to_entries[] | {key, type: (.value | type)}]'
      
      # Reformat JSONL with consistent key ordering
      jq -cS '.' file.jsonl > normalized.jsonl
      ```
      
      ### Debugging jq Expressions
      
      ```bash
      # Use debug to print intermediate values to stderr
      jq '.items[] | debug | select(.active)' file.json
      
      # Use @text to see what jq thinks a value is
      jq '.field | @text' file.json
      
      # Use type to check value types
      jq '.field | type' file.json
      
      # Build expressions incrementally
      jq '.' file.json                    # Start: see full structure
      jq '.items' file.json               # Navigate to array
      jq '.items[]' file.json             # Iterate array
      jq '.items[] | .name' file.json     # Extract field
      jq '.items[] | select(.active)' file.json  # Filter
      
      # Common error: "Cannot iterate over null"
      # Fix: use ? operator
      jq '.items[]?' file.json            # Won't error if items is null
      
      # Common error: "null is not iterable"
      # Fix: default empty array
      jq '(.items // [])[]' file.json
      ```
      
      ### Integration with Other Tools
      
      ```bash
      # rg prefilter then jq parse
      rg '"error"' app.jsonl | jq -r '.message'
      
      # jq output to column for alignment
      jq -r '[.name, .status, .duration] | @tsv' app.jsonl | column -t
      
      # jq output to sort/uniq for frequency
      jq -r '.error_type' errors.jsonl | sort | uniq -c | sort -rn
      
      # jq to CSV for spreadsheet import
      jq -r '[.timestamp, .level, .message] | @csv' app.jsonl > export.csv
      
      # jq with xargs for per-line processing
      jq -r '.file_path' manifest.jsonl | xargs wc -l
      ```
      
      ---
      
      ## lnav -- Log File Navigator
      
      Interactive terminal-based log viewer with SQL support, automatic format detection, timeline view, and filtering. Ideal for exploratory analysis.
      
      ### Installation
      
      ```bash
      # macOS
      brew install lnav
      
      # Ubuntu/Debian
      sudo apt install lnav
      
      # Windows (via Chocolatey)
      choco install lnav
      
      # From source
      curl -LO https://github.com/tstack/lnav/releases/download/v0.12.2/lnav-0.12.2-linux-musl-x86_64.zip
      unzip lnav-0.12.2-linux-musl-x86_64.zip
      sudo cp lnav-0.12.2/lnav /usr/local/bin/
      
      # Verify
      lnav -V
      ```
      
      ### Key Features
      
      | Feature | Access | Description |
      |---------|--------|-------------|
      | Auto-detect format | Automatic | Recognizes syslog, Apache, nginx, JSON, and many more |
      | SQL queries | `:` then SQL | Run SQL against log data |
      | Filter in/out | `i` / `o` | Interactive include/exclude filters |
      | Bookmarks | `m` | Mark lines for later reference |
      | Timeline | `t` | Show time histogram |
      | Pretty print | `p` | Toggle pretty-printing JSON |
      | Headless mode | `-n -c "..."` | Non-interactive command execution |
      | Compressed files | Automatic | Handles .gz, .bz2, .xz transparently |
      
      ### Essential Commands
      
      ```bash
      # Open log file(s)
      lnav app.log
      lnav /var/log/syslog /var/log/auth.log   # multiple files, merged by timestamp
      
      # Open JSONL logs
      lnav app.jsonl
      
      # Open compressed logs
      lnav app.log.gz
      
      # Open all logs in a directory
      lnav /var/log/myapp/
      
      # Headless mode: run query and output results
      lnav -n -c ";SELECT count(*) FROM logline WHERE log_level = 'error'" app.log
      
      # Headless mode: filter and export
      lnav -n -c ";SELECT log_time, log_body FROM logline WHERE log_level = 'error'" \
        -c ":write-csv-to errors.csv" app.log
      
      # Headless mode: get stats
      lnav -n -c ";SELECT log_level, count(*) as cnt FROM logline GROUP BY log_level ORDER BY cnt DESC" app.log
      ```
      
      ### SQL Mode Recipes
      
      ```sql
      -- Error count by hour
      SELECT strftime('%Y-%m-%d %H', log_time) as hour, count(*) as errors
      FROM logline WHERE log_level = 'error'
      GROUP BY hour ORDER BY hour;
      
      -- Top error messages
      SELECT log_body, count(*) as cnt
      FROM logline WHERE log_level = 'error'
      GROUP BY log_body ORDER BY cnt DESC LIMIT 10;
      
      -- Time between events
      SELECT log_time, log_body,
        julianday(log_time) - julianday(lag(log_time) OVER (ORDER BY log_time)) as gap_days
      FROM logline WHERE log_level = 'error';
      
      -- Log volume over time
      SELECT strftime('%Y-%m-%d %H:%M', log_time) as minute, count(*) as lines
      FROM logline
      GROUP BY minute ORDER BY minute;
      ```
      
      ### Custom Log Formats
      
      ```json
      // ~/.lnav/formats/installed/myapp.json
      {
        "myapp_log": {
          "title": "My Application Log",
          "regex": {
            "std": {
              "pattern": "^(?<timestamp>\\d{4}-\\d{2}-\\d{2}T\\d{2}:\\d{2}:\\d{2})\\s+\\[(?<level>\\w+)\\]\\s+(?<body>.*)"
            }
          },
          "timestamp-format": ["%Y-%m-%dT%H:%M:%S"],
          "level": {
            "error": "ERROR",
            "warning": "WARN",
            "info": "INFO",
            "debug": "DEBUG"
          }
        }
      }
      ```
      
      ### Interactive Keyboard Shortcuts
      
      | Key | Action |
      |-----|--------|
      | `/` | Search forward (regex) |
      | `n` / `N` | Next/previous search match |
      | `i` | Toggle filter: show only matching lines |
      | `o` | Toggle filter: hide matching lines |
      | `TAB` | Switch between views (log, text, help) |
      | `t` | Toggle timeline histogram |
      | `m` | Set bookmark on current line |
      | `u` / `U` | Next/previous bookmark |
      | `z` / `Z` | Zoom in/out on timeline |
      | `p` | Toggle pretty-print for JSON |
      | `e` / `E` | Next/previous error |
      | `w` / `W` | Next/previous warning |
      | `:` | Enter command mode |
      | `;` | Enter SQL query mode |
      
      ---
      
      ## angle-grinder (agrind) -- Log Pipeline Aggregation
      
      Pipeline-based aggregation tool designed for log analysis. Think SQL-like queries in a streaming pipeline syntax.
      
      ### Installation
      
      ```bash
      # Via cargo (all platforms)
      cargo install ag
      
      # macOS
      brew install angle-grinder
      
      # Verify
      agrind --version
      ```
      
      ### Pipeline Syntax
      
      ```
      <input_pattern> | <operator1> | <operator2> | ...
      ```
      
      ### Essential Commands
      
      ```bash
      # Count log levels
      cat app.log | agrind '* | parse "* [*] *" as ts, level, msg | count by level'
      
      # Top URLs
      cat access.log | agrind '* | parse "* * * * * * *" as ip, _, _, ts, method, url, status | count by url | sort by _count desc | head 10'
      
      # Average response time by endpoint
      cat access.log | agrind '* | parse "* *ms" as prefix, duration | avg of duration by prefix'
      
      # Error frequency over time
      cat app.log | agrind '* | parse "*T*:*:* [ERROR]*" as date, hour, min, sec, msg | count by hour'
      
      # Filter then aggregate
      cat app.log | agrind '* | where level == "error" | count by msg | sort by _count desc'
      
      # JSON log fields
      cat app.jsonl | agrind '* | json | where level == "error" | count by message'
      ```
      
      ### Operators Reference
      
      | Operator | Purpose | Example |
      |----------|---------|---------|
      | `parse` | Extract fields with pattern | `parse "* [*] *" as a, b, c` |
      | `json` | Parse JSON log lines | `json` |
      | `where` | Filter rows | `where level == "error"` |
      | `count` | Count (optionally by group) | `count by level` |
      | `sum` | Sum a field | `sum of bytes` |
      | `avg` | Average a field | `avg of duration` |
      | `min` / `max` | Min/max of field | `min of response_time` |
      | `sort` | Sort results | `sort by _count desc` |
      | `head` | Limit results | `head 10` |
      | `uniq` | Unique values | `uniq by user_id` |
      | `percentile` | Percentile calc | `p50 of duration, p99 of duration` |
      
      ---
      
      ## rg (ripgrep) -- Fast Pattern Search
      
      Already covered extensively in file-search skill. Here are log-specific flags and patterns.
      
      ### Log-Specific Flags
      
      | Flag | Purpose | Example |
      |------|---------|---------|
      | `-c` | Count matches per file | `rg -c "ERROR" /var/log/*.log` |
      | `-l` | List files with matches | `rg -l "timeout" /var/log/` |
      | `-L` | List files without matches | `rg -L "healthy" /var/log/` |
      | `--stats` | Show match statistics | `rg --stats "error" app.log` |
      | `-A N` | Show N lines after match | `rg -A5 "Exception" app.log` |
      | `-B N` | Show N lines before match | `rg -B3 "FATAL" app.log` |
      | `-C N` | Show N lines context | `rg -C5 "crash" app.log` |
      | `-U` | Multiline matching | `rg -U "Error.*\n.*at " app.log` |
      | `--json` | JSON output format | `rg --json "error" app.log` |
      | `-a` | Search binary files | `rg -a "pattern" binary.log` |
      | `--line-buffered` | Flush per line (for tail) | `tail -f app.log \| rg --line-buffered "error"` |
      | `-F` | Fixed string (no regex) | `rg -F "stack[0]" app.log` |
      | `-f FILE` | Patterns from file | `rg -f patterns.txt app.log` |
      | `-v` | Invert match | `rg -v "DEBUG" app.log` |
      
      ### Log Search Recipes
      
      ```bash
      # Find errors across all log files recursively
      rg "ERROR|FATAL|CRITICAL" /var/log/
      
      # Count errors per file, sorted
      rg -c "ERROR" /var/log/ 2>/dev/null | sort -t: -k2 -rn
      
      # Find stack traces (multiline)
      rg -U "Exception.*\n(\s+at .*\n)+" app.log
      
      # Extract timestamps of errors
      rg "ERROR" app.log | rg -o "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}"
      
      # Search compressed log files
      rg -z "error" app.log.gz
      
      # Search JSONL for specific field value (text-level, fast but approximate)
      rg '"level":"error"' app.jsonl
      
      # Search JSONL for value in specific key (avoid matching wrong key)
      rg '"user_id":"user-42"' app.jsonl
      
      # Negative lookahead: errors that are NOT timeouts
      rg "ERROR(?!.*timeout)" app.log
      
      # Time-bounded search (extract lines between two timestamps)
      rg "2026-03-08T1[4-5]:" app.log
      ```
      
      ### rg JSON Output Mode
      
      ```bash
      # Get structured output from rg (useful for programmatic processing)
      rg --json "error" app.log | jq -c 'select(.type == "match") | {file: .data.path.text, line: .data.line_number, text: .data.lines.text}'
      
      # Count matches with file info
      rg --json "error" app.log | jq -c 'select(.type == "summary") | .data.stats'
      ```
      
      ---
      
      ## awk -- Column-Based Log Processing
      
      Pre-installed on all Unix systems. Best for space/tab delimited logs with consistent column structure.
      
      ### Common Recipes
      
      ```bash
      # Apache/nginx combined log format columns:
      # $1=IP $2=ident $3=user $4=date $5=time $6=tz $7=method $8=path $9=proto $10=status $11=size
      
      # Status code distribution
      awk '{print $9}' access.log | sort | uniq -c | sort -rn
      
      # Requests per IP
      awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -20
      
      # 5xx errors with paths
      awk '$9 >= 500 {print $1, $9, $7}' access.log
      
      # Average response size
      awk '{sum += $10; n++} END {printf "Avg: %.0f bytes\n", sum/n}' access.log
      
      # Requests per minute
      awk '{print substr($4, 2, 17)}' access.log | sort | uniq -c | tail -20
      
      # Bandwidth by path
      awk '{bytes[$7] += $10} END {for (p in bytes) printf "%10d %s\n", bytes[p], p}' access.log | sort -rn | head -20
      
      # Custom delimiter (e.g., pipe-separated)
      awk -F'|' '{print $3, $5}' custom.log
      
      # Time-range filter (syslog format)
      awk '/^Mar  8 14:/ {print}' syslog
      
      # Calculate time difference between first and last line
      awk 'NR==1 {first=$1" "$2} END {last=$1" "$2; print "From:", first, "To:", last}' app.log
      ```
      
      ### awk for Key-Value Logs
      
      ```bash
      # Parse key=value format (logfmt)
      awk '{
        for (i=1; i<=NF; i++) {
          split($i, kv, "=")
          if (kv[1] == "duration") sum += kv[2]; n++
        }
      } END {print "avg_duration=" sum/n}' app.log
      
      # Extract specific key from logfmt
      awk '{
        for (i=1; i<=NF; i++) {
          split($i, kv, "=")
          if (kv[1] == "status" && kv[2] >= 500) print $0
        }
      }' app.log
      ```
      
      ---
      
      ## GNU parallel -- Parallel Log Processing
      
      Splits work across CPU cores for processing large log files.
      
      ### Installation
      
      ```bash
      # macOS
      brew install parallel
      
      # Ubuntu/Debian
      sudo apt install parallel
      
      # Verify
      parallel --version
      ```
      
      ### Essential Commands
      
      ```bash
      # Process multiple log files in parallel
      ls /var/log/app-*.jsonl | parallel "jq -c 'select(.level == \"error\")' {} > {.}_errors.jsonl"
      
      # Split large file and process chunks in parallel
      split -l 200000 huge.jsonl /tmp/chunk_
      ls /tmp/chunk_* | parallel "jq -r '.message' {} | sort | uniq -c" | sort -rn | head -20
      
      # Parallel grep across many files
      fd -e jsonl . /var/log/ | parallel "rg -c '\"error\"' {} 2>/dev/null" | sort -t: -k2 -rn
      
      # Pipe-based parallelism (no temp files)
      parallel --pipe -L 50000 "jq -c 'select(.level == \"error\")'" < huge.jsonl > errors.jsonl
      
      # Parallel with progress bar
      ls /var/log/app-*.jsonl | parallel --bar "jq -sc 'length' {}" | awk '{sum+=$1} END {print sum, "total lines"}'
      
      # Number of jobs (default: CPU cores)
      ls *.jsonl | parallel -j 4 "jq -c 'select(.status >= 500)' {}"
      ```
      
      ### Combining with split
      
      ```bash
      # Full workflow: split, process in parallel, merge results
      FILE=huge.jsonl
      CHUNKS=/tmp/log_chunks
      mkdir -p "$CHUNKS"
      
      # Split
      split -l 100000 "$FILE" "$CHUNKS/chunk_"
      
      # Process in parallel
      ls "$CHUNKS"/chunk_* | parallel "jq -r 'select(.level == \"error\") | .message' {}" |
        sort | uniq -c | sort -rn > error_summary.txt
      
      # Cleanup
      rm -rf "$CHUNKS"
      ```
      
      ---
      
      ## Miller (mlr) -- CSV/TSV Log Analysis
      
      Like awk, sed, and jq combined but specifically for structured record data (CSV, TSV, JSON).
      
      ### Installation
      
      ```bash
      # macOS
      brew install miller
      
      # Ubuntu/Debian
      sudo apt install miller
      
      # Windows
      choco install miller
      
      # Verify
      mlr --version
      ```
      
      ### Essential Commands
      
      ```bash
      # View CSV with headers
      mlr --csv head -n 10 access_log.csv
      
      # Filter rows
      mlr --csv filter '$status >= 400' access_log.csv
      
      # Sort by column
      mlr --csv sort-by -nr duration access_log.csv
      
      # Statistics
      mlr --csv stats1 -a min,max,mean,p95 -f duration access_log.csv
      
      # Group by with stats
      mlr --csv stats1 -a count,mean -f duration -g endpoint access_log.csv
      
      # Convert formats
      mlr --c2j cat access_log.csv          # CSV to JSON
      mlr --c2t cat access_log.csv          # CSV to TSV (table)
      mlr --j2c cat access_log.json         # JSON to CSV
      mlr --c2p cat access_log.csv          # CSV to pretty-print table
      
      # Top-N by group
      mlr --csv top -n 5 -f duration -g endpoint access_log.csv
      
      # Add computed fields
      mlr --csv put '$error = ($status >= 400 ? "yes" : "no")' access_log.csv
      
      # Decimate (sample every Nth row)
      mlr --csv sample -k 100 huge_log.csv
      
      # Uniq count
      mlr --csv count-distinct -f status access_log.csv
      
      # Histogram
      mlr --csv decimate -g status -n 1 access_log.csv | mlr --csv count-distinct -f status
      
      # Join two CSV files
      mlr --csv join -j user_id -f users.csv then sort-by user_id access_log.csv
      ```
      
      ### TSV from jq to mlr Pipeline
      
      ```bash
      # Extract JSONL to TSV, then use mlr for analysis
      jq -r '[.timestamp, .level, .duration_ms, .path] | @tsv' app.jsonl > /tmp/extracted.tsv
      mlr --tsvlite --from /tmp/extracted.tsv \
        label timestamp,level,duration_ms,path then \
        filter '$level == "error"' then \
        stats1 -a count,mean -f duration_ms -g path then \
        sort-by -nr count
      ```
      
      ---
      
      ## Tool Integration Cheat Sheet
      
      ### Combining Tools
      
      ```bash
      # fd + rg + jq: find files, prefilter, extract
      fd -e jsonl . logs/ | xargs rg -l '"error"' | xargs jq -c 'select(.level == "error") | {ts: .timestamp, msg: .message}'
      
      # rg + jq + column: search, extract, format
      rg '"timeout"' app.jsonl | jq -r '[.timestamp, .service, .message] | @tsv' | column -t
      
      # jq + sort + uniq: aggregate without slurp
      jq -r '.error_type' errors.jsonl | sort | uniq -c | sort -rn
      
      # tail + rg + jq: live monitoring with extraction
      tail -f app.jsonl | rg --line-buffered '"error"' | jq --unbuffered -r '[.timestamp, .message] | @tsv'
      
      # fd + parallel + jq: parallel extraction across many files
      fd -e jsonl . logs/ | parallel "jq -c 'select(.level == \"error\")' {}" > all_errors.jsonl
      
      # jq + mlr: structured extraction then statistical analysis
      jq -r '[.path, .duration_ms, .status] | @csv' app.jsonl | \
        mlr --csv label path,duration,status then \
        stats1 -a p50,p95,p99 -f duration -g path then \
        sort-by -nr p95
      
      # lnav + headless SQL: non-interactive queries
      lnav -n -c ";SELECT log_level, count(*) FROM logline GROUP BY log_level" app.log
      ```
      
      ### Decision Guide: Which Combination?
      
      ```
      Task: Explore unknown log file
        --> lnav (interactive, auto-detects format)
      
      Task: Quick search for pattern
        --> rg "pattern" file.log
      
      Task: Extract fields from JSONL
        --> jq -r '[.field1, .field2] | @tsv' file.jsonl
      
      Task: Count/aggregate JSONL (<100MB)
        --> jq -sc 'group_by(.x) | map(...)' file.jsonl
      
      Task: Count/aggregate JSONL (>100MB)
        --> jq -r '.field' file.jsonl | sort | uniq -c | sort -rn
      
      Task: Search large JSONL then extract
        --> rg "pattern" file.jsonl | jq -r '.field'
      
      Task: CSV/TSV log statistics
        --> mlr --csv stats1 -a mean,p95 -f duration file.csv
      
      Task: Process many log files in parallel
        --> fd -e jsonl . dir/ | parallel "jq ..."
      
      Task: Pipeline aggregation on text logs
        --> cat file.log | agrind '* | parse ... | count by ...'
      
      Task: Live monitoring with filtering
        --> tail -f file.jsonl | rg --line-buffered "x" | jq --unbuffered '.'
      ```
      
  • scripts
    • .gitkeep 0 B · in bundle
  • SKILL.md 15.8 KB
    ---
    name: log-ops
    description: "Log analysis and JSONL processing - structured extraction, cross-log correlation, timeline reconstruction, pattern search"
    license: MIT
    allowed-tools: "Read Edit Write Bash Glob Grep Agent"
    metadata:
      author: claude-mods
      related-skills: data-processing, debug-ops, monitoring-ops, file-search, introspect
    ---
    
    # Log Operations
    
    Practical patterns for analyzing log files -- especially JSONL format used in agent conversation logs, benchmark outputs, and structured application logs.
    
    ## Log Format Decision Tree
    
    ```
    Unknown Log File
    │
    ├─ Is it one JSON object per line?
    │  ├─ Yes ──────────────────────── JSONL
    │  │  ├─ Small file (<100MB)
    │  │  │  └─ jq for extraction, jq -s for aggregation
    │  │  ├─ Large file (100MB-1GB)
    │  │  │  └─ rg prefilter then pipe to jq
    │  │  └─ Huge file (>1GB)
    │  │     └─ split + parallel jq, or jq --stream
    │  │
    │  └─ No
    │     ├─ Is it one large JSON object/array?
    │     │  └─ Yes ──────────────── Single JSON
    │     │     └─ jq --stream for SAX-style, or jq directly if fits in memory
    │     │
    │     ├─ Does it have key=value pairs?
    │     │  └─ Yes ──────────────── Structured (logfmt / key-value)
    │     │     └─ rg for search, awk/sd for extraction, angle-grinder for aggregation
    │     │
    │     ├─ Does it follow syslog format? (timestamp hostname service[pid]: message)
    │     │  └─ Yes ──────────────── Syslog
    │     │     └─ rg for search, awk for column extraction, lnav for interactive
    │     │
    │     ├─ Is it space/tab delimited with consistent columns?
    │     │  └─ Yes ──────────────── Column-based (access logs, CSV)
    │     │     └─ awk for extraction, mlr for CSV, rg for pattern search
    │     │
    │     └─ Mixed or unstructured
    │        └─ Plain text ─────────── Freeform
    │           └─ rg for search, rg -A/-B for context, lnav for exploration
    ```
    
    ## Prerequisites
    
    **Required** (must be installed):
    - `rg` (ripgrep) - text search, prefiltering. Install: `cargo install ripgrep` / `choco install ripgrep`
    - `jq` - JSON/JSONL extraction and transformation. Install: `brew install jq` / `choco install jq`
    
    **Optional** (enhanced capabilities, gracefully degraded without):
    - `lnav` - interactive log exploration with SQL queries. Install: `brew install lnav` / WSL: `apt install lnav`
    - `agrind` (angle-grinder) - pipeline aggregation syntax. Install: `cargo install ag`
    - `mlr` (Miller) - CSV/TSV log analysis. Install: `brew install miller` / `choco install miller`
    - `GNU parallel` - parallel processing of split files. Install: `brew install parallel`
    
    > All patterns in this skill work with just rg + jq. Optional tools add interactive exploration (lnav), pipeline aggregation (agrind), and tabular analysis (mlr).
    
    ## Tool Selection Matrix
    
    | Tool | Best For | Speed | Required? |
    |------|----------|-------|-----------|
    | `rg` (ripgrep) | Raw pattern matching in any format | Fastest | Yes |
    | `jq` | JSONL structured extraction and transformation | Fast | Yes |
    | `jq -s` | JSONL aggregation (slurp all lines into array) | Medium (loads all into memory) | Yes (part of jq) |
    | `lnav` | Interactive exploration, SQL over logs | Interactive | Optional |
    | `agrind` (angle-grinder) | Pipeline aggregation and counting | Fast | Optional |
    | `awk` | Column-based log formats, field extraction | Fast | Pre-installed |
    | `mlr` (Miller) | CSV/TSV log analysis, statistics | Fast | Optional |
    | `fd` + `rg` | Searching across many log directories | Fast | Pre-installed in dev-shell |
    | `GNU parallel` | Splitting large files for parallel processing | N/A (orchestrator) | Optional |
    
    ### When to Use What
    
    ```
    Need to...
    │
    ├─ Find lines matching a pattern
    │  └─ rg (always fastest for text search)
    │
    ├─ Extract specific fields from JSONL
    │  └─ jq -r '[.field1, .field2] | @tsv'
    │
    ├─ Count/aggregate over JSONL
    │  └─ jq -sc 'group_by(.field) | map({key: .[0].field, n: length})'
    │
    ├─ Search JSONL by value then format results
    │  └─ rg '"error"' file.jsonl | jq -r '.message'  (two-stage)
    │
    ├─ Explore interactively with filtering/SQL
    │  └─ lnav file.log
    │
    ├─ Aggregate with pipeline syntax
    │  └─ agrind '* | parse "* * *" as ts, level, msg | count by level'
    │
    ├─ Extract columns from space-delimited logs
    │  └─ awk '{print $1, $4, $7}' access.log
    │
    └─ Process CSV/TSV logs with headers
       └─ mlr --csv filter '$status >= 400' then stats1 -a count -f status
    ```
    
    ## JSONL Quick Reference
    
    The most common format for structured logs. One JSON object per line, no trailing commas, no wrapping array.
    
    ### Stream Filtering (line by line, constant memory)
    
    ```bash
    # Filter by field value
    jq -c 'select(.level == "error")' app.jsonl
    
    # Filter by nested field
    jq -c 'select(.request.method == "POST")' app.jsonl
    
    # Filter by multiple conditions
    jq -c 'select(.level == "error" and .status >= 500)' app.jsonl
    
    # Filter by array contains
    jq -c 'select(.tags | index("critical"))' app.jsonl
    
    # Filter by field existence
    jq -c 'select(.stack_trace != null)' app.jsonl
    
    # Negate a filter
    jq -c 'select(.level != "debug")' app.jsonl
    ```
    
    ### Field Extraction
    
    ```bash
    # Extract single field
    jq -r '.message' app.jsonl
    
    # Extract multiple fields as TSV
    jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl
    
    # Extract with default for missing fields
    jq -r '.error_code // "none"' app.jsonl
    
    # Extract nested field safely
    jq -r '.response.headers["content-type"] // "unknown"' app.jsonl
    ```
    
    ### Aggregation (requires slurp: loads entire file)
    
    ```bash
    # Count by field value
    jq -sc 'group_by(.level) | map({level: .[0].level, count: length})' app.jsonl
    
    # Top-N most common values
    jq -sc '[.[].error_type] | group_by(.) | map({type: .[0], count: length}) | sort_by(-.count) | .[:10]' app.jsonl
    
    # Sum a numeric field
    jq -sc 'map(.duration_ms) | add' app.jsonl
    
    # Average
    jq -sc 'map(.duration_ms) | add / length' app.jsonl
    
    # Min and max
    jq -sc 'map(.duration_ms) | {min: min, max: max}' app.jsonl
    ```
    
    ### Nested Extraction (agent logs, complex structures)
    
    ```bash
    # Extract tool calls from conversation logs
    jq -c '.content[]? | select(.type == "tool_use") | .name' conversation.jsonl
    
    # De-escape nested JSON strings
    jq -c '.content | fromjson' app.jsonl
    
    # Flatten nested arrays
    jq -c '[.events[]? | .action]' app.jsonl
    
    # Extract from arrays of objects
    jq -c '.results[]? | select(.passed == false) | {test: .name, error: .message}' results.jsonl
    ```
    
    ### Two-Stage Pipeline (rg for speed, jq for structure)
    
    ```bash
    # Fast prefilter then structured extraction
    rg '"error"' app.jsonl | jq -r '[.timestamp, .message] | @tsv'
    
    # Search for specific value then aggregate
    rg '"timeout"' app.jsonl | jq -sc 'length'
    
    # Pattern match then extract
    rg '"user_id":"u-123"' app.jsonl | jq -c '{ts: .timestamp, action: .action}'
    ```
    
    ### Time-Range Filtering
    
    ```bash
    # Filter by timestamp range (ISO 8601 string comparison works)
    jq -c 'select(.timestamp > "2026-03-08T10:00" and .timestamp < "2026-03-08T11:00")' app.jsonl
    
    # Events in the last N minutes (using epoch seconds)
    jq -c --arg cutoff "$(date -d '30 minutes ago' +%s)" 'select((.timestamp | sub("\\.[0-9]+Z$"; "Z") | fromdate) > ($cutoff | tonumber))' app.jsonl
    
    # Extract hour for histogram
    jq -r '.timestamp | split("T")[1] | split(":")[0]' app.jsonl | sort | uniq -c
    ```
    
    ### Cross-File Join
    
    ```bash
    # Extract IDs from one file, search in another
    jq -r '.request_id' errors.jsonl | while read id; do
      rg "\"$id\"" responses.jsonl | jq -c '{id: .request_id, status: .status}'
    done
    
    # Faster: build lookup, then join
    jq -r '.request_id' errors.jsonl | sort -u > /tmp/error_ids.txt
    rg -Ff /tmp/error_ids.txt responses.jsonl | jq -c '{id: .request_id, status: .status}'
    
    # Join two JSONL files by key using jq --slurpfile
    jq --slurpfile lookup <(jq -sc 'map({(.id): .}) | add' lookup.jsonl) \
      '. + ($lookup[0][.ref_id] // {})' main.jsonl
    ```
    
    ## Plain Text Log Patterns
    
    ### Pattern Search with Context
    
    ```bash
    # Show 5 lines before and after each match
    rg -B5 -A5 "OutOfMemoryError" app.log
    
    # Show only matching files
    rg -l "FATAL" /var/log/
    
    # Count matches per file
    rg -c "ERROR" /var/log/*.log | sort -t: -k2 -rn
    
    # Multiline patterns (stack traces)
    rg -U "Exception.*\n(\s+at .*\n)+" app.log
    ```
    
    ### Column Extraction with awk
    
    ```bash
    # Apache/nginx access log: extract status codes
    awk '{print $9}' access.log | sort | uniq -c | sort -rn
    
    # Extract specific time range from syslog
    awk '$0 >= "Mar  8 10:00" && $0 <= "Mar  8 11:00"' syslog
    
    # Calculate average response time (column 11)
    awk '{sum += $11; n++} END {print sum/n}' access.log
    
    # Filter by status code and show URL + response time
    awk '$9 >= 500 {print $7, $11"ms"}' access.log
    ```
    
    ### Live Monitoring
    
    ```bash
    # Follow with filtering
    tail -f app.log | rg --line-buffered "ERROR"
    
    # Follow JSONL and extract fields
    tail -f app.jsonl | jq --unbuffered -r '[.timestamp, .level, .message] | @tsv'
    
    # Follow multiple files
    tail -f /var/log/service-*.log | rg --line-buffered "error|warn"
    ```
    
    ## Timeline Reconstruction
    
    ### Extracting and Sorting by Timestamp
    
    ```bash
    # Merge multiple log files by timestamp
    sort -t' ' -k1,2 service-a.log service-b.log > timeline.log
    
    # JSONL: sort by timestamp field
    jq -sc 'sort_by(.timestamp)[]' combined.jsonl > sorted.jsonl
    
    # Extract timestamps and calculate gaps
    jq -r '.timestamp' app.jsonl | awk '
      NR > 1 {
        cmd = "date -d \"" prev "\" +%s"; cmd | getline t1; close(cmd)
        cmd = "date -d \"" $0 "\" +%s"; cmd | getline t2; close(cmd)
        gap = t2 - t1
        if (gap > 5) print gap "s gap before " $0
      }
      { prev = $0 }
    '
    
    # Quick duration between first and last event
    jq -sc '{start: .[0].timestamp, end: .[-1].timestamp}' app.jsonl
    ```
    
    ### Calculating Durations Between Events
    
    ```bash
    # Duration between paired events (start/end)
    jq -sc '
      group_by(.request_id) |
      map(
        (map(select(.event == "start")) | .[0].timestamp) as $start |
        (map(select(.event == "end")) | .[0].timestamp) as $end |
        {id: .[0].request_id, start: $start, end: $end}
      )
    ' events.jsonl
    
    # Identify the slowest phase
    jq -sc '
      sort_by(.timestamp) |
      [range(1; length) | {
        from: .[.-1].event,
        to: .[.].event,
        gap: ((.[.].ts_epoch) - (.[.-1].ts_epoch))
      }] |
      sort_by(-.gap) | .[0]
    ' events.jsonl
    ```
    
    ## Cross-Log Correlation
    
    ### By Correlation ID
    
    ```bash
    # Find a request across all service logs
    fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {}
    
    # Build a timeline for a single request
    fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {} \; | jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .event] | @tsv'
    ```
    
    ### By Timestamp Window
    
    ```bash
    # Find events within 2 seconds of a known event
    # First get the target timestamp
    TARGET="2026-03-08T14:23:15"
    jq -c --arg t "$TARGET" '
      select(
        .timestamp > ($t | sub("15$"; "13")) and
        .timestamp < ($t | sub("15$"; "17"))
      )
    ' other-service.jsonl
    ```
    
    ### By Session/User
    
    ```bash
    # Reconstruct a user session across log files
    fd -e jsonl . /var/log/ -x rg "\"user-42\"" {} \; |
      jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .action] | @tsv'
    ```
    
    ## Large File Strategies
    
    ### Search Recent Only
    
    ```bash
    # Last 10,000 lines (fast for append-only logs)
    tail -n 10000 huge.log | rg "pattern"
    
    # Last N lines of JSONL with structured extraction
    tail -n 5000 huge.jsonl | jq -c 'select(.level == "error")'
    ```
    
    ### Split for Parallel Processing
    
    ```bash
    # Split into 100K-line chunks
    split -l 100000 huge.jsonl /tmp/chunk_
    
    # Process in parallel
    fd 'chunk_' /tmp/ -x jq -c 'select(.level == "error")' {} > errors.jsonl
    
    # With GNU parallel
    split -l 100000 huge.jsonl /tmp/chunk_
    ls /tmp/chunk_* | parallel 'jq -c "select(.level == \"error\")" {} >> /tmp/errors.jsonl'
    ```
    
    ### Streaming for Huge Single JSON
    
    ```bash
    # SAX-style processing of a huge JSON array
    jq --stream 'select(.[0][0] == "results" and .[0][-1] == "status") | .[1]' huge.json
    
    # Extract items from a huge array without loading all
    jq -cn --stream 'fromstream(1 | truncate_stream(inputs))' huge-array.json
    ```
    
    ### Two-Stage Always
    
    ```bash
    # ALWAYS faster: rg filters text, jq parses survivors
    rg '"error"' huge.jsonl | jq -r '.message'
    
    # vs. SLOW: jq reads and parses every line
    jq -r 'select(.level == "error") | .message' huge.jsonl
    ```
    
    ## Search Across Directories
    
    ### Multi-Directory Patterns
    
    ```bash
    # Find all JSONL files with errors across trial directories
    fd -e jsonl . trials/ -x rg -l '"error"' {}
    
    # Count errors per log file across directories
    fd -e jsonl . trials/ -x bash -c 'echo "$(rg -c "\"error\"" "$1" 2>/dev/null || echo 0) $1"' _ {}
    
    # Extract and aggregate across directories
    fd -e jsonl . trials/ -x jq -c 'select(.level == "error") | {file: input_filename, msg: .message}' {}
    
    # Build summary table from multiple runs
    for dir in trials/*/; do
      total=$(wc -l < "$dir/results.jsonl")
      errors=$(rg -c '"error"' "$dir/results.jsonl" 2>/dev/null || echo 0)
      echo -e "$dir\t$total\t$errors"
    done | column -t -N DIR,TOTAL,ERRORS
    ```
    
    ## Common Gotchas
    
    | Gotcha | Why It Hurts | Fix |
    |--------|-------------|-----|
    | `jq -s` on huge files loads everything into memory | OOM crash or swap thrashing on files over ~500MB | Use streaming: `rg` prefilter, `jq --stream`, or `split` + parallel |
    | JSONL with embedded newlines in string values | Line-by-line tools (rg, awk, head) split a single record across lines | Use `jq -c` to re-compact, or `jq -R 'fromjson?'` to skip malformed lines |
    | rg matches JSON keys, not just values | `rg "error"` matches `{"error_count": 0}` which is not an error | Use `rg '"level":"error"'` or pipe to `jq 'select(.level == "error")'` |
    | Timezone mismatches in timestamp comparisons | Events appear out of order or time ranges miss data | Normalize to UTC before comparing: `jq '.timestamp |= sub("\\+.*"; "Z")'` |
    | Unicode and escape sequences in log messages | jq chokes on invalid UTF-8 or double-escaped strings | Prefilter with `rg -a` (binary mode), or use `jq -R` for raw strings |
    | Inconsistent JSON schemas across log lines | `jq` errors on lines missing expected fields | Use `//` operator for defaults: `.field // "missing"` and `?` for optional: `.arr[]?` |
    | Forgetting `-c` flag with jq on JSONL | jq pretty-prints each line, output is no longer valid JSONL | Always use `jq -c` when output feeds into another JSONL consumer |
    | tail -f with jq buffering | Output appears delayed or not at all | Use `jq --unbuffered` or `stdbuf -oL jq` |
    | Sorting JSONL by timestamp without slurp | `sort` command does lexicographic sort on whole lines, not by field | Either `jq -sc 'sort_by(.timestamp)[]'` or extract timestamp prefix first |
    | Assuming log files are complete | Logs may be rotated, compressed, or still being written | Check for `.gz` rotated files: `fd -e gz . /var/log/ -x zcat {} \| rg pattern` |
    | Single quotes in jq on Windows | PowerShell/cmd do not handle single quotes the same as bash | Use double quotes with escaped inner quotes, or write jq filter to a file |
    
    ## Reference Files
    
    | File | Contents | Lines |
    |------|----------|-------|
    | `references/jsonl-patterns.md` | JSONL extraction, aggregation, transformation, comparison, and performance patterns | ~700 |
    | `references/analysis-workflows.md` | Agent conversation analysis, application log analysis, benchmark result parsing, cross-directory workflows | ~600 |
    | `references/tool-setup.md` | Installation and configuration for jq, lnav, angle-grinder, rg, awk, GNU parallel, Miller | ~450 |
    
    ## See Also
    
    - **data-processing** -- JSON/YAML/TOML processing with jq and yq
    - **debug-ops** -- Systematic debugging methodology, log-based debugging section
    - **monitoring-ops** -- Production observability, alerting, dashboards
    - **file-search** -- Finding files with fd, searching code with rg
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related