log-ops
Log analysis and JSONL processing - structured extraction, cross-log correlation, timeline reconstruction, pattern search
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/log-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Log Operations
Practical patterns for analyzing log files -- especially JSONL format used in agent conversation logs, benchmark outputs, and structured application logs.
Log Format Decision Tree
Unknown Log File
│
├─ Is it one JSON object per line?
│ ├─ Yes ──────────────────────── JSONL
│ │ ├─ Small file (<100MB)
│ │ │ └─ jq for extraction, jq -s for aggregation
│ │ ├─ Large file (100MB-1GB)
│ │ │ └─ rg prefilter then pipe to jq
│ │ └─ Huge file (>1GB)
│ │ └─ split + parallel jq, or jq --stream
│ │
│ └─ No
│ ├─ Is it one large JSON object/array?
│ │ └─ Yes ──────────────── Single JSON
│ │ └─ jq --stream for SAX-style, or jq directly if fits in memory
│ │
│ ├─ Does it have key=value pairs?
│ │ └─ Yes ──────────────── Structured (logfmt / key-value)
│ │ └─ rg for search, awk/sd for extraction, angle-grinder for aggregation
│ │
│ ├─ Does it follow syslog format? (timestamp hostname service[pid]: message)
│ │ └─ Yes ──────────────── Syslog
│ │ └─ rg for search, awk for column extraction, lnav for interactive
│ │
│ ├─ Is it space/tab delimited with consistent columns?
│ │ └─ Yes ──────────────── Column-based (access logs, CSV)
│ │ └─ awk for extraction, mlr for CSV, rg for pattern search
│ │
│ └─ Mixed or unstructured
│ └─ Plain text ─────────── Freeform
│ └─ rg for search, rg -A/-B for context, lnav for exploration
Prerequisites
Required (must be installed):
rg(ripgrep) - text search, prefiltering. Install:cargo install ripgrep/choco install ripgrepjq- JSON/JSONL extraction and transformation. Install:brew install jq/choco install jq
Optional (enhanced capabilities, gracefully degraded without):
lnav- interactive log exploration with SQL queries. Install:brew install lnav/ WSL:apt install lnavagrind(angle-grinder) - pipeline aggregation syntax. Install:cargo install agmlr(Miller) - CSV/TSV log analysis. Install:brew install miller/choco install millerGNU parallel- parallel processing of split files. Install:brew install parallel
All patterns in this skill work with just rg + jq. Optional tools add interactive exploration (lnav), pipeline aggregation (agrind), and tabular analysis (mlr).
Tool Selection Matrix
| Tool | Best For | Speed | Required? |
|---|---|---|---|
rg (ripgrep) |
Raw pattern matching in any format | Fastest | Yes |
jq |
JSONL structured extraction and transformation | Fast | Yes |
jq -s |
JSONL aggregation (slurp all lines into array) | Medium (loads all into memory) | Yes (part of jq) |
lnav |
Interactive exploration, SQL over logs | Interactive | Optional |
agrind (angle-grinder) |
Pipeline aggregation and counting | Fast | Optional |
awk |
Column-based log formats, field extraction | Fast | Pre-installed |
mlr (Miller) |
CSV/TSV log analysis, statistics | Fast | Optional |
fd + rg |
Searching across many log directories | Fast | Pre-installed in dev-shell |
GNU parallel |
Splitting large files for parallel processing | N/A (orchestrator) | Optional |
When to Use What
Need to...
│
├─ Find lines matching a pattern
│ └─ rg (always fastest for text search)
│
├─ Extract specific fields from JSONL
│ └─ jq -r '[.field1, .field2] | @tsv'
│
├─ Count/aggregate over JSONL
│ └─ jq -sc 'group_by(.field) | map({key: .[0].field, n: length})'
│
├─ Search JSONL by value then format results
│ └─ rg '"error"' file.jsonl | jq -r '.message' (two-stage)
│
├─ Explore interactively with filtering/SQL
│ └─ lnav file.log
│
├─ Aggregate with pipeline syntax
│ └─ agrind '* | parse "* * *" as ts, level, msg | count by level'
│
├─ Extract columns from space-delimited logs
│ └─ awk '{print $1, $4, $7}' access.log
│
└─ Process CSV/TSV logs with headers
└─ mlr --csv filter '$status >= 400' then stats1 -a count -f status
JSONL Quick Reference
The most common format for structured logs. One JSON object per line, no trailing commas, no wrapping array.
Stream Filtering (line by line, constant memory)
# Filter by field value
jq -c 'select(.level == "error")' app.jsonl
# Filter by nested field
jq -c 'select(.request.method == "POST")' app.jsonl
# Filter by multiple conditions
jq -c 'select(.level == "error" and .status >= 500)' app.jsonl
# Filter by array contains
jq -c 'select(.tags | index("critical"))' app.jsonl
# Filter by field existence
jq -c 'select(.stack_trace != null)' app.jsonl
# Negate a filter
jq -c 'select(.level != "debug")' app.jsonl
Field Extraction
# Extract single field
jq -r '.message' app.jsonl
# Extract multiple fields as TSV
jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl
# Extract with default for missing fields
jq -r '.error_code // "none"' app.jsonl
# Extract nested field safely
jq -r '.response.headers["content-type"] // "unknown"' app.jsonl
Aggregation (requires slurp: loads entire file)
# Count by field value
jq -sc 'group_by(.level) | map({level: .[0].level, count: length})' app.jsonl
# Top-N most common values
jq -sc '[.[].error_type] | group_by(.) | map({type: .[0], count: length}) | sort_by(-.count) | .[:10]' app.jsonl
# Sum a numeric field
jq -sc 'map(.duration_ms) | add' app.jsonl
# Average
jq -sc 'map(.duration_ms) | add / length' app.jsonl
# Min and max
jq -sc 'map(.duration_ms) | {min: min, max: max}' app.jsonl
Nested Extraction (agent logs, complex structures)
# Extract tool calls from conversation logs
jq -c '.content[]? | select(.type == "tool_use") | .name' conversation.jsonl
# De-escape nested JSON strings
jq -c '.content | fromjson' app.jsonl
# Flatten nested arrays
jq -c '[.events[]? | .action]' app.jsonl
# Extract from arrays of objects
jq -c '.results[]? | select(.passed == false) | {test: .name, error: .message}' results.jsonl
Two-Stage Pipeline (rg for speed, jq for structure)
# Fast prefilter then structured extraction
rg '"error"' app.jsonl | jq -r '[.timestamp, .message] | @tsv'
# Search for specific value then aggregate
rg '"timeout"' app.jsonl | jq -sc 'length'
# Pattern match then extract
rg '"user_id":"u-123"' app.jsonl | jq -c '{ts: .timestamp, action: .action}'
Time-Range Filtering
# Filter by timestamp range (ISO 8601 string comparison works)
jq -c 'select(.timestamp > "2026-03-08T10:00" and .timestamp < "2026-03-08T11:00")' app.jsonl
# Events in the last N minutes (using epoch seconds)
jq -c --arg cutoff "$(date -d '30 minutes ago' +%s)" 'select((.timestamp | sub("\\.[0-9]+Z$"; "Z") | fromdate) > ($cutoff | tonumber))' app.jsonl
# Extract hour for histogram
jq -r '.timestamp | split("T")[1] | split(":")[0]' app.jsonl | sort | uniq -c
Cross-File Join
# Extract IDs from one file, search in another
jq -r '.request_id' errors.jsonl | while read id; do
rg "\"$id\"" responses.jsonl | jq -c '{id: .request_id, status: .status}'
done
# Faster: build lookup, then join
jq -r '.request_id' errors.jsonl | sort -u > /tmp/error_ids.txt
rg -Ff /tmp/error_ids.txt responses.jsonl | jq -c '{id: .request_id, status: .status}'
# Join two JSONL files by key using jq --slurpfile
jq --slurpfile lookup <(jq -sc 'map({(.id): .}) | add' lookup.jsonl) \
'. + ($lookup[0][.ref_id] // {})' main.jsonl
Plain Text Log Patterns
Pattern Search with Context
# Show 5 lines before and after each match
rg -B5 -A5 "OutOfMemoryError" app.log
# Show only matching files
rg -l "FATAL" /var/log/
# Count matches per file
rg -c "ERROR" /var/log/*.log | sort -t: -k2 -rn
# Multiline patterns (stack traces)
rg -U "Exception.*\n(\s+at .*\n)+" app.log
Column Extraction with awk
# Apache/nginx access log: extract status codes
awk '{print $9}' access.log | sort | uniq -c | sort -rn
# Extract specific time range from syslog
awk '$0 >= "Mar 8 10:00" && $0 <= "Mar 8 11:00"' syslog
# Calculate average response time (column 11)
awk '{sum += $11; n++} END {print sum/n}' access.log
# Filter by status code and show URL + response time
awk '$9 >= 500 {print $7, $11"ms"}' access.log
Live Monitoring
# Follow with filtering
tail -f app.log | rg --line-buffered "ERROR"
# Follow JSONL and extract fields
tail -f app.jsonl | jq --unbuffered -r '[.timestamp, .level, .message] | @tsv'
# Follow multiple files
tail -f /var/log/service-*.log | rg --line-buffered "error|warn"
Timeline Reconstruction
Extracting and Sorting by Timestamp
# Merge multiple log files by timestamp
sort -t' ' -k1,2 service-a.log service-b.log > timeline.log
# JSONL: sort by timestamp field
jq -sc 'sort_by(.timestamp)[]' combined.jsonl > sorted.jsonl
# Extract timestamps and calculate gaps
jq -r '.timestamp' app.jsonl | awk '
NR > 1 {
cmd = "date -d \"" prev "\" +%s"; cmd | getline t1; close(cmd)
cmd = "date -d \"" $0 "\" +%s"; cmd | getline t2; close(cmd)
gap = t2 - t1
if (gap > 5) print gap "s gap before " $0
}
{ prev = $0 }
'
# Quick duration between first and last event
jq -sc '{start: .[0].timestamp, end: .[-1].timestamp}' app.jsonl
Calculating Durations Between Events
# Duration between paired events (start/end)
jq -sc '
group_by(.request_id) |
map(
(map(select(.event == "start")) | .[0].timestamp) as $start |
(map(select(.event == "end")) | .[0].timestamp) as $end |
{id: .[0].request_id, start: $start, end: $end}
)
' events.jsonl
# Identify the slowest phase
jq -sc '
sort_by(.timestamp) |
[range(1; length) | {
from: .[.-1].event,
to: .[.].event,
gap: ((.[.].ts_epoch) - (.[.-1].ts_epoch))
}] |
sort_by(-.gap) | .[0]
' events.jsonl
Cross-Log Correlation
By Correlation ID
# Find a request across all service logs
fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {}
# Build a timeline for a single request
fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {} \; | jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .event] | @tsv'
By Timestamp Window
# Find events within 2 seconds of a known event
# First get the target timestamp
TARGET="2026-03-08T14:23:15"
jq -c --arg t "$TARGET" '
select(
.timestamp > ($t | sub("15$"; "13")) and
.timestamp < ($t | sub("15$"; "17"))
)
' other-service.jsonl
By Session/User
# Reconstruct a user session across log files
fd -e jsonl . /var/log/ -x rg "\"user-42\"" {} \; |
jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .action] | @tsv'
Large File Strategies
Search Recent Only
# Last 10,000 lines (fast for append-only logs)
tail -n 10000 huge.log | rg "pattern"
# Last N lines of JSONL with structured extraction
tail -n 5000 huge.jsonl | jq -c 'select(.level == "error")'
Split for Parallel Processing
# Split into 100K-line chunks
split -l 100000 huge.jsonl /tmp/chunk_
# Process in parallel
fd 'chunk_' /tmp/ -x jq -c 'select(.level == "error")' {} > errors.jsonl
# With GNU parallel
split -l 100000 huge.jsonl /tmp/chunk_
ls /tmp/chunk_* | parallel 'jq -c "select(.level == \"error\")" {} >> /tmp/errors.jsonl'
Streaming for Huge Single JSON
# SAX-style processing of a huge JSON array
jq --stream 'select(.[0][0] == "results" and .[0][-1] == "status") | .[1]' huge.json
# Extract items from a huge array without loading all
jq -cn --stream 'fromstream(1 | truncate_stream(inputs))' huge-array.json
Two-Stage Always
# ALWAYS faster: rg filters text, jq parses survivors
rg '"error"' huge.jsonl | jq -r '.message'
# vs. SLOW: jq reads and parses every line
jq -r 'select(.level == "error") | .message' huge.jsonl
Search Across Directories
Multi-Directory Patterns
# Find all JSONL files with errors across trial directories
fd -e jsonl . trials/ -x rg -l '"error"' {}
# Count errors per log file across directories
fd -e jsonl . trials/ -x bash -c 'echo "$(rg -c "\"error\"" "$1" 2>/dev/null || echo 0) $1"' _ {}
# Extract and aggregate across directories
fd -e jsonl . trials/ -x jq -c 'select(.level == "error") | {file: input_filename, msg: .message}' {}
# Build summary table from multiple runs
for dir in trials/*/; do
total=$(wc -l < "$dir/results.jsonl")
errors=$(rg -c '"error"' "$dir/results.jsonl" 2>/dev/null || echo 0)
echo -e "$dir\t$total\t$errors"
done | column -t -N DIR,TOTAL,ERRORS
Common Gotchas
| Gotcha | Why It Hurts | Fix |
|---|---|---|
jq -s on huge files loads everything into memory |
OOM crash or swap thrashing on files over ~500MB | Use streaming: rg prefilter, jq --stream, or split + parallel |
| JSONL with embedded newlines in string values | Line-by-line tools (rg, awk, head) split a single record across lines | Use jq -c to re-compact, or jq -R 'fromjson?' to skip malformed lines |
| rg matches JSON keys, not just values | rg "error" matches {"error_count": 0} which is not an error |
Use rg '"level":"error"' or pipe to jq 'select(.level == "error")' |
| Timezone mismatches in timestamp comparisons | Events appear out of order or time ranges miss data | Normalize to UTC before comparing: jq '.timestamp |= sub("\\+.*"; "Z")' |
| Unicode and escape sequences in log messages | jq chokes on invalid UTF-8 or double-escaped strings | Prefilter with rg -a (binary mode), or use jq -R for raw strings |
| Inconsistent JSON schemas across log lines | jq errors on lines missing expected fields |
Use // operator for defaults: .field // "missing" and ? for optional: .arr[]? |
Forgetting -c flag with jq on JSONL |
jq pretty-prints each line, output is no longer valid JSONL | Always use jq -c when output feeds into another JSONL consumer |
| tail -f with jq buffering | Output appears delayed or not at all | Use jq --unbuffered or stdbuf -oL jq |
| Sorting JSONL by timestamp without slurp | sort command does lexicographic sort on whole lines, not by field |
Either jq -sc 'sort_by(.timestamp)[]' or extract timestamp prefix first |
| Assuming log files are complete | Logs may be rotated, compressed, or still being written | Check for .gz rotated files: fd -e gz . /var/log/ -x zcat {} \| rg pattern |
| Single quotes in jq on Windows | PowerShell/cmd do not handle single quotes the same as bash | Use double quotes with escaped inner quotes, or write jq filter to a file |
Reference Files
| File | Contents | Lines |
|---|---|---|
references/jsonl-patterns.md |
JSONL extraction, aggregation, transformation, comparison, and performance patterns | ~700 |
references/analysis-workflows.md |
Agent conversation analysis, application log analysis, benchmark result parsing, cross-directory workflows | ~600 |
references/tool-setup.md |
Installation and configuration for jq, lnav, angle-grinder, rg, awk, GNU parallel, Miller | ~450 |
See Also
- data-processing -- JSON/YAML/TOML processing with jq and yq
- debug-ops -- Systematic debugging methodology, log-based debugging section
- monitoring-ops -- Production observability, alerting, dashboards
- file-search -- Finding files with fd, searching code with rg
Files (claude-mods)
-
assets
-
.gitkeep 0 B · in bundle
-
-
references
-
analysis-workflows.md 17.9 KB
# Analysis Workflows Reference Practical end-to-end workflows for common log analysis tasks. Each workflow is self-contained with commands you can copy and adapt. --- ## Agent Conversation Log Analysis Claude Code and other AI agents produce JSONL conversation logs with nested content blocks. These workflows extract actionable information from those logs. ### Extract All Tool Calls ```bash # List every tool call in chronological order jq -c ' select(.role == "assistant") | .content[]? | select(.type == "tool_use") | {tool: .name, id: .id} ' conversation.jsonl # Count tool usage frequency jq -r ' select(.role == "assistant") | .content[]? | select(.type == "tool_use") | .name ' conversation.jsonl | sort | uniq -c | sort -rn # Extract tool calls with their inputs (summarized) jq -c ' select(.role == "assistant") | .content[]? | select(.type == "tool_use") | { tool: .name, input_preview: (.input | tostring | .[:100]) } ' conversation.jsonl ``` ### Identify What Code Was Written ```bash # Find all Write tool calls and extract file paths jq -r ' select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Write") | .input.file_path ' conversation.jsonl # Find all Edit tool calls with file paths and old/new strings jq -c ' select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Edit") | {file: .input.file_path, old: (.input.old_string | .[:60]), new: (.input.new_string | .[:60])} ' conversation.jsonl # Find all Bash commands that were run jq -r ' select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Bash") | .input.command ' conversation.jsonl # Files created vs modified echo "=== Files Created (Write) ===" jq -r 'select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Write") | .input.file_path' conversation.jsonl | sort -u echo "=== Files Modified (Edit) ===" jq -r 'select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Edit") | .input.file_path' conversation.jsonl | sort -u ``` ### Find Error Messages and Repeated Attempts ```bash # Find tool results that indicate errors jq -c ' select(.role == "tool") | .content[]? | select(.type == "text") | select(.text | test("error|Error|ERROR|failed|Failed|FAILED|exception|Exception")) | {text_preview: (.text | .[:200])} ' conversation.jsonl # Find retry patterns (same tool called multiple times with similar input) jq -r ' select(.role == "assistant") | .content[]? | select(.type == "tool_use") | "\(.name)\t\(.input | tostring | .[:80])" ' conversation.jsonl | sort | uniq -c | sort -rn | head -20 # Count consecutive failures (same tool, error in result) jq -sc ' [to_entries[] | select(.value.role == "tool") | {idx: .key, has_error: (.value.content | tostring | test("error|Error|failed|Failed"))} ] | map(select(.has_error)) | length ' conversation.jsonl ``` ### Calculate Phase Timings ```bash # If messages have timestamps, calculate time between phases jq -sc ' map(select(.timestamp != null)) | sort_by(.timestamp) | . as $msgs | { total_messages: length, first: .[0].timestamp, last: .[-1].timestamp, tool_calls: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use")] | length, reading_ops: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use" and (.name == "Read" or .name == "Glob" or .name == "Grep"))] | length, writing_ops: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use" and (.name == "Write" or .name == "Edit"))] | length, bash_ops: [.[] | select(.role == "assistant") | .content[]? | select(.type == "tool_use" and .name == "Bash")] | length } ' conversation.jsonl # Phase breakdown by sequential grouping jq -c ' select(.role == "assistant") | .content[]? | select(.type == "tool_use") | if .name == "Read" or .name == "Glob" or .name == "Grep" then "READING" elif .name == "Write" or .name == "Edit" then "WRITING" elif .name == "Bash" then "EXECUTING" else "OTHER" end ' conversation.jsonl | uniq -c ``` ### Extract Thinking Blocks and Reasoning ```bash # Extract thinking/reasoning content jq -r ' select(.role == "assistant") | .content[]? | select(.type == "thinking") | .thinking ' conversation.jsonl # Extract text responses (non-tool, non-thinking) jq -r ' select(.role == "assistant") | .content[]? | select(.type == "text") | .text ' conversation.jsonl # Summary of assistant responses jq -c ' select(.role == "assistant") | { has_thinking: (.content | any(.type == "thinking")), has_text: (.content | any(.type == "text")), tool_calls: [.content[]? | select(.type == "tool_use") | .name] } ' conversation.jsonl ``` ### Build a Timeline of Actions ```bash # Full action timeline jq -r ' if .role == "user" then "USER: " + (.content | if type == "string" then .[:100] else (.[] | select(.type == "text") | .text | .[:100]) end) elif .role == "assistant" then (.content[]? | if .type == "tool_use" then "TOOL: " + .name + " " + (.input | tostring | .[:80]) elif .type == "text" then "TEXT: " + (.text | .[:100]) else empty end ) elif .role == "tool" then "RESULT: " + (.content | tostring | .[:100]) else empty end ' conversation.jsonl # Condensed timeline (just tool calls and results) jq -c ' if .role == "assistant" then .content[]? | select(.type == "tool_use") | {action: "call", tool: .name} elif .role == "tool" then {action: "result", success: (.content | tostring | test("error|Error|failed") | not)} else empty end ' conversation.jsonl ``` --- ## Application Log Analysis ### Error Rate Over Time ```bash # Errors per minute jq -r 'select(.level == "error") | .timestamp | .[:16]' app.jsonl | sort | uniq -c # Errors per hour with total context jq -rsc ' group_by(.timestamp | .[:13]) | map({ hour: .[0].timestamp | .[:13], total: length, errors: (map(select(.level == "error")) | length) }) | map("\(.hour)\t\(.total)\t\(.errors)\t\(.errors * 100 / .total | round)%") | .[] ' app.jsonl | column -t -N HOUR,TOTAL,ERRORS,RATE # Error rate spike detection (>2x average) jq -sc ' group_by(.timestamp | .[:13]) | map({hour: .[0].timestamp | .[:13], errors: (map(select(.level == "error")) | length)}) | (map(.errors) | add / length) as $avg | map(select(.errors > ($avg * 2))) | map("\(.hour): \(.errors) errors (avg: \($avg | round))")[] ' app.jsonl ``` ### Slow Request Identification ```bash # Top 20 slowest requests jq -sc ' sort_by(-.duration_ms) | .[:20] | .[] | [.timestamp, .method, .path, "\(.duration_ms)ms"] | @tsv ' app.jsonl | column -t # Slow requests by endpoint (p95) jq -sc ' group_by(.path) | map({ path: .[0].path, count: length, p50: (map(.duration_ms) | sort | .[length * 0.5 | floor]), p95: (map(.duration_ms) | sort | .[length * 0.95 | floor]), p99: (map(.duration_ms) | sort | .[length * 0.99 | floor]) }) | sort_by(-.p95) | .[:10] ' app.jsonl # Requests exceeding SLA (e.g., 500ms) jq -c 'select(.duration_ms > 500) | {path, duration_ms, timestamp}' app.jsonl | jq -rsc 'group_by(.path) | map({path: .[0].path, count: length, worst: (map(.duration_ms) | max)}) | sort_by(-.count) | .[] | [.path, .count, .worst] | @tsv' | column -t -N ENDPOINT,SLA_VIOLATIONS,WORST_MS ``` ### Error Correlation ```bash # Which errors occur together in the same time window? jq -rsc ' map(select(.level == "error")) | group_by(.timestamp | .[:16]) | map(select(length > 1)) | map([.[].message] | unique | sort) | group_by(.) | map({errors: .[0], co_occurrences: length}) | sort_by(-.co_occurrences) | .[:10] ' app.jsonl # Errors that always precede another error jq -rsc ' map(select(.level == "error")) | sort_by(.timestamp) | [range(1; length) | {before: .[. - 1].message, after: .[.].message}] | group_by([.before, .after]) | map({sequence: .[0], count: length}) | sort_by(-.count) | .[:10] ' app.jsonl # Error clusters (errors within 5 seconds of each other) jq -rsc ' map(select(.level == "error")) | sort_by(.timestamp) | . as $errs | [range(1; length) | select( (($errs[.].ts_epoch // 0) - ($errs[. - 1].ts_epoch // 0)) < 5 ) | {ts: $errs[.].timestamp, msg: $errs[.].message} ] ' app.jsonl ``` ### User Session Reconstruction ```bash # Reconstruct a single user session jq -c 'select(.user_id == "user-42")' app.jsonl | jq -sc 'sort_by(.timestamp) | .[] | [.timestamp, .action, .path // .event] | @tsv' | column -t # Session summary for all users jq -sc ' group_by(.user_id) | map({ user: .[0].user_id, events: length, first_seen: (sort_by(.timestamp) | .[0].timestamp), last_seen: (sort_by(.timestamp) | .[-1].timestamp), unique_actions: ([.[].action] | unique | length), errors: (map(select(.level == "error")) | length) }) | sort_by(-.events) ' app.jsonl # User journey (sequence of page views) jq -r 'select(.user_id == "user-42" and .event == "page_view") | .path' app.jsonl ``` ### Deployment Impact Analysis ```bash # Compare error rates before and after deployment DEPLOY_TIME="2026-03-08T14:30:00" echo "=== Before Deployment ===" jq -c --arg t "$DEPLOY_TIME" 'select(.timestamp < $t)' app.jsonl | jq -sc '{total: length, errors: (map(select(.level == "error")) | length)}' echo "=== After Deployment ===" jq -c --arg t "$DEPLOY_TIME" 'select(.timestamp >= $t)' app.jsonl | jq -sc '{total: length, errors: (map(select(.level == "error")) | length)}' # New error types after deployment BEFORE=$(jq -r --arg t "$DEPLOY_TIME" 'select(.timestamp < $t and .level == "error") | .message' app.jsonl | sort -u) AFTER=$(jq -r --arg t "$DEPLOY_TIME" 'select(.timestamp >= $t and .level == "error") | .message' app.jsonl | sort -u) comm -13 <(echo "$BEFORE") <(echo "$AFTER") # Response time comparison echo "=== Response Times Before ===" jq -sc --arg t "$DEPLOY_TIME" ' map(select(.timestamp < $t and .duration_ms != null)) | {avg: (map(.duration_ms) | add / length | round), p95: (map(.duration_ms) | sort | .[length * 0.95 | floor])} ' app.jsonl echo "=== Response Times After ===" jq -sc --arg t "$DEPLOY_TIME" ' map(select(.timestamp >= $t and .duration_ms != null)) | {avg: (map(.duration_ms) | add / length | round), p95: (map(.duration_ms) | sort | .[length * 0.95 | floor])} ' app.jsonl ``` --- ## Benchmark and Test Result Analysis ### Parse Structured Test Results ```bash # CTRF JSON format (Common Test Report Format) jq -r '.results.tests[] | select(.status == "failed") | [.name, .message // "no message"] | @tsv' ctrf-report.json # CTRF summary jq '{ total: .results.summary.tests, passed: .results.summary.passed, failed: .results.summary.failed, skipped: .results.summary.skipped, duration: "\(.results.summary.duration)ms" }' ctrf-report.json # JUnit XML (convert to JSON first with xq or yq) yq -p xml '.testsuites.testsuite.testcase[] | select(.failure != null) | ."+@name"' junit-results.xml # TAP (Test Anything Protocol) - extract failures rg "^not ok" test-output.tap | sd 'not ok \d+ - ' '' ``` ### Compare Pass/Fail Rates Across Runs ```bash # Compare multiple CTRF reports for report in results/*/ctrf-report.json; do dir=$(dirname "$report" | xargs basename) passed=$(jq '.results.summary.passed' "$report") failed=$(jq '.results.summary.failed' "$report") total=$(jq '.results.summary.tests' "$report") echo -e "$dir\t$passed\t$failed\t$total" done | column -t -N RUN,PASSED,FAILED,TOTAL # Find tests that regressed (passed before, fail now) jq -r '.results.tests[] | select(.status == "passed") | .name' run1/ctrf-report.json | sort > /tmp/passed_before.txt jq -r '.results.tests[] | select(.status == "failed") | .name' run2/ctrf-report.json | sort > /tmp/failed_after.txt comm -12 /tmp/passed_before.txt /tmp/failed_after.txt # Flaky test detection (tests that flip between runs) for report in results/*/ctrf-report.json; do jq -r '.results.tests[] | "\(.name)\t\(.status)"' "$report" done | sort | awk -F'\t' ' {status[$1] = status[$1] " " $2} END { for (test in status) { if (status[test] ~ /passed/ && status[test] ~ /failed/) { print "FLAKY:", test, status[test] } } } ' ``` ### Performance Regression Detection ```bash # Compare timing data between runs jq -sc '[.[] | {name: .name, duration: .duration}]' run1/results.jsonl > /tmp/run1_times.json jq -sc '[.[] | {name: .name, duration: .duration}]' run2/results.jsonl > /tmp/run2_times.json # Find tests that got significantly slower (>20% regression) jq -sc ' [., input] | (.[0] | map({(.name): .duration}) | add) as $before | (.[1] | map({(.name): .duration}) | add) as $after | [$before | keys[] | select($after[.] != null) | { name: ., before: $before[.], after: $after[.], change_pct: (($after[.] - $before[.]) / $before[.] * 100 | round) } | select(.change_pct > 20) ] | sort_by(-.change_pct) ' /tmp/run1_times.json /tmp/run2_times.json # Aggregate metrics across trial directories for dir in trials/trial-*/; do trial=$(basename "$dir") if [ -f "$dir/metrics.jsonl" ]; then avg=$(jq -sc 'map(.duration) | add / length | round' "$dir/metrics.jsonl") p95=$(jq -sc 'map(.duration) | sort | .[length * 0.95 | floor]' "$dir/metrics.jsonl") echo -e "$trial\t$avg\t$p95" fi done | column -t -N TRIAL,AVG_MS,P95_MS ``` ### Aggregate Metrics Across Trial Directories ```bash # Build summary from multiple benchmark runs fd -t d 'trial-' trials/ -x bash -c ' trial=$(basename "$1") if [ -f "$1/results.jsonl" ]; then total=$(wc -l < "$1/results.jsonl") passed=$(jq -c "select(.passed == true)" "$1/results.jsonl" | wc -l) failed=$((total - passed)) echo -e "$trial\t$total\t$passed\t$failed" fi ' _ {} | sort | column -t -N TRIAL,TOTAL,PASSED,FAILED # Combine all results into one file with trial label fd -t d 'trial-' trials/ -x bash -c ' trial=$(basename "$1") jq -c --arg trial "$trial" ". + {trial: \$trial}" "$1/results.jsonl" ' _ {} > combined_results.jsonl # Then aggregate across all trials jq -sc ' group_by(.trial) | map({ trial: .[0].trial, total: length, pass_rate: ((map(select(.passed == true)) | length) / length * 100 | round), avg_duration: (map(.duration) | add / length | round) }) | sort_by(.trial) ' combined_results.jsonl ``` --- ## Cross-Directory Analysis ### Search Pattern Across All Log Directories ```bash # Find which log files contain a specific error fd -e jsonl -e log . /var/log/services/ -x rg -l "ConnectionTimeout" {} # Count occurrences per directory fd -e jsonl . logs/ -x bash -c ' count=$(rg -c "error" "$1" 2>/dev/null || echo 0) echo -e "$(dirname "$1" | xargs basename)\t$(basename "$1")\t$count" ' _ {} | sort -t$'\t' -k3 -rn | column -t -N DIR,FILE,ERRORS # Search for a pattern and show matching lines with source file fd -e jsonl . logs/ -x bash -c ' rg "\"error\"" "$1" 2>/dev/null | while read line; do echo "$1: $line" done ' _ {} ``` ### Build Summary Table from Multiple Log Files ```bash # Summary statistics per log file echo -e "FILE\tLINES\tERRORS\tWARNS\tFIRST_TS\tLAST_TS" fd -e jsonl . logs/ | while read f; do lines=$(wc -l < "$f") errors=$(rg -c '"error"' "$f" 2>/dev/null || echo 0) warns=$(rg -c '"warn"' "$f" 2>/dev/null || echo 0) first=$(head -1 "$f" | jq -r '.timestamp // "unknown"') last=$(tail -1 "$f" | jq -r '.timestamp // "unknown"') echo -e "$(basename "$f")\t$lines\t$errors\t$warns\t$first\t$last" done | column -t # Health check across all services fd -e jsonl -d 1 . /var/log/services/ -x bash -c ' svc=$(basename "$1" .jsonl) last_error=$(tac "$1" | jq -r "select(.level == \"error\") | .timestamp" 2>/dev/null | head -1) error_count=$(rg -c "\"error\"" "$1" 2>/dev/null || echo 0) echo -e "$svc\t$error_count\t${last_error:-none}" ' _ {} | sort | column -t -N SERVICE,ERRORS,LAST_ERROR ``` ### Identify Common Failure Patterns Across Runs ```bash # Extract all error messages across trial directories fd -e jsonl . trials/ -x jq -r 'select(.level == "error") | .message' {} | sort | uniq -c | sort -rn | head -20 # Find which trials share the same failure fd -e jsonl . trials/ -x bash -c ' trial=$(echo "$1" | rg -o "trial-[^/]+") jq -r "select(.level == \"error\") | .message" "$1" 2>/dev/null | while read msg; do echo -e "$trial\t$msg"; done ' _ {} | sort -t$'\t' -k2 | awk -F'\t' ' prev != $2 { if (NR > 1 && count > 1) print count, prev_msg, trials; count=0; trials="" } { count++; trials = trials " " $1; prev = $2; prev_msg = $2 } END { if (count > 1) print count, prev_msg, trials } ' | sort -rn | head -10 # Correlation: which errors appear together fd -e jsonl . trials/ -x bash -c ' trial=$(echo "$1" | rg -o "trial-[^/]+") errors=$(jq -r "select(.level == \"error\") | .message" "$1" 2>/dev/null | sort -u | paste -sd "|") [ -n "$errors" ] && echo -e "$trial\t$errors" ' _ {} | sort -t$'\t' -k2 | uniq -f1 -c | sort -rn ``` ### fd + rg + jq Composition ```bash # The canonical three-stage pipeline for multi-directory log analysis: # 1. fd: find the files # 2. rg: prefilter for speed # 3. jq: structured extraction # Example: find all timeout errors across services, extract details fd -e jsonl . /var/log/ | # find log files xargs rg -l '"timeout"' | # filter to files with timeouts xargs -I{} jq -c ' select(.message | test("timeout")) | {file: input_filename, ts: .timestamp, svc: .service, msg: .message} ' {} # Example: aggregate error counts by service across all log directories fd -e jsonl . /var/log/ -x rg -c '"error"' {} | # count errors per file awk -F: '{ split($1, parts, "/") svc = parts[length(parts)-1] gsub(/\.jsonl$/, "", svc) sum[svc] += $2 } END { for (s in sum) print sum[s], s }' | sort -rn # Example: find the most recent error across all services fd -e jsonl . /var/log/ -x tail -1 {} | # last line of each file jq -sc ' map(select(.level == "error")) | sort_by(.timestamp) | .[-1] | {service: .service, timestamp: .timestamp, message: .message} ' ``` -
jsonl-patterns.md 17 KB
# JSONL Patterns Reference Comprehensive patterns for working with JSONL (JSON Lines) files -- one JSON object per line, the dominant format for structured logs, agent conversation records, and streaming data. ## JSONL Basics ### Format Rules - One valid JSON object per line - No trailing commas between lines - No wrapping array or outer object - Each line is independently parseable - Newlines within string values must be escaped as `\n` ### Streaming vs Slurp ```bash # STREAMING (default): processes one line at a time, constant memory jq -c 'select(.level == "error")' app.jsonl # SLURP (-s): loads ALL lines into a single array, requires memory for entire file jq -sc 'group_by(.level)' app.jsonl # Rule of thumb: # File < 100MB --> slurp is fine # File 100MB-1GB --> slurp with caution, prefer streaming + sort/uniq # File > 1GB --> never slurp, use streaming or split+parallel ``` ### Key jq Flags for JSONL | Flag | Purpose | Example | |------|---------|---------| | `-c` | Compact output (one line per object) | `jq -c '.' file.jsonl` | | `-r` | Raw string output (no quotes) | `jq -r '.message' file.jsonl` | | `-s` | Slurp all lines into array | `jq -s 'length' file.jsonl` | | `-e` | Exit with error if output is false/null | `jq -e '.status == 200' line.json` | | `-R` | Read each line as raw string | `jq -R 'fromjson? // empty' messy.jsonl` | | `--stream` | SAX-style path/value pairs | `jq --stream '.' huge.json` | | `--slurpfile` | Load a file as variable | `jq --slurpfile ids ids.json 'select(.id | IN($ids[][]))' data.jsonl` | | `--arg` | Pass string variable | `jq --arg name "foo" 'select(.name == $name)' data.jsonl` | | `--argjson` | Pass JSON variable | `jq --argjson min 100 'select(.count > $min)' data.jsonl` | | `--unbuffered` | Flush output after each line | `tail -f app.jsonl \| jq --unbuffered -r '.message'` | --- ## Extraction Patterns ### Select by Field Value ```bash # Exact match jq -c 'select(.level == "error")' app.jsonl # Numeric comparison jq -c 'select(.status >= 400)' app.jsonl # String contains jq -c 'select(.message | test("timeout"))' app.jsonl # Regex match jq -c 'select(.path | test("^/api/v[0-9]+/users"))' app.jsonl # Case-insensitive match jq -c 'select(.message | test("error"; "i"))' app.jsonl # Null check jq -c 'select(.error != null)' app.jsonl # Boolean field jq -c 'select(.retry == true)' app.jsonl ``` ### Select by Nested Field ```bash # Dot notation for nesting jq -c 'select(.request.method == "POST")' app.jsonl # Deep nesting jq -c 'select(.context.user.role == "admin")' app.jsonl # Safe navigation (no error if path missing) jq -c 'select(.request?.headers?["authorization"] != null)' app.jsonl ``` ### Select by Array Contains ```bash # Array contains value jq -c 'select(.tags | index("critical"))' app.jsonl # Any element matches condition jq -c 'select(.events | any(.type == "error"))' app.jsonl # All elements match condition jq -c 'select(.checks | all(.passed == true))' app.jsonl # Array length jq -c 'select((.retries | length) > 3)' app.jsonl ``` ### Extract and Flatten Nested Structures ```bash # Flatten one level of nesting jq -c '{timestamp, level, msg: .message, user: .context.user.id}' app.jsonl # Explode array into separate lines jq -c '.events[]' app.jsonl # Flatten array with parent context jq -c '. as $parent | .events[] | {request_id: $parent.request_id, event: .type, ts: .timestamp}' app.jsonl # Extract from array of objects jq -c '.results[] | select(.score < 0.5) | {name, score}' results.jsonl # Recursive descent (find all values for a key at any depth) jq -c '.. | .error_message? // empty' app.jsonl ``` ### Handle Optional and Nullable Fields ```bash # Default value for missing field jq -r '.region // "unknown"' app.jsonl # Default for nested missing field jq -r '.response.body.error // .response.status_text // "no error info"' app.jsonl # Skip lines where field is missing (instead of outputting null) jq -r '.optional_field // empty' app.jsonl # Coalesce multiple possible fields jq -r '(.error_message // .err_msg // .error // "none")' app.jsonl # Type check before access jq -c 'if .data | type == "array" then .data | length else 0 end' app.jsonl ``` ### Multi-Level Nesting (Agent Conversation Logs) ```bash # Claude Code conversation logs have deeply nested tool calls # Structure: {role, content: [{type: "tool_use", name, input}, ...]} # Extract all tool call names jq -c '.content[]? | select(.type == "tool_use") | .name' conversation.jsonl # Extract tool inputs jq -c '.content[]? | select(.type == "tool_use") | {tool: .name, input: .input}' conversation.jsonl # Extract text content blocks jq -r '.content[]? | select(.type == "text") | .text' conversation.jsonl # Extract tool results jq -c '.content[]? | select(.type == "tool_result") | {tool_use_id, content}' conversation.jsonl # Find tool calls that contain specific patterns in their input jq -c '.content[]? | select(.type == "tool_use" and (.input | tostring | test("SELECT")))' conversation.jsonl ``` ### De-Escape Nested JSON Strings ```bash # When a field contains a JSON string that needs parsing jq -c '.payload | fromjson' app.jsonl # Safe de-escape (skip if not valid JSON) jq -c '.payload | fromjson? // {raw: .}' app.jsonl # Double-escaped JSON (escaped twice) jq -c '.data | fromjson | fromjson' app.jsonl # Extract field from de-escaped nested JSON jq -r '.payload | fromjson | .result.status' app.jsonl # Handle mixed escaped/unescaped jq -c 'if (.payload | type) == "string" then .payload | fromjson else .payload end' app.jsonl ``` --- ## Aggregation Patterns All aggregation patterns use `-s` (slurp) which loads the entire file into memory. For large files, prefilter with `rg` first. ### Count by Field Value ```bash # Count per level jq -sc 'group_by(.level) | map({level: .[0].level, count: length})' app.jsonl # Count per status code jq -sc 'group_by(.status) | map({status: .[0].status, count: length}) | sort_by(-.count)' app.jsonl # Count unique values jq -sc '[.[].user_id] | unique | length' app.jsonl # Frequency distribution jq -rsc 'group_by(.level) | map("\(.[0].level)\t\(length)") | .[]' app.jsonl ``` ### Sum, Average, Min, Max ```bash # Sum jq -sc 'map(.bytes) | add' app.jsonl # Average jq -sc 'map(.duration_ms) | add / length' app.jsonl # Min and max jq -sc 'map(.duration_ms) | {min: min, max: max, avg: (add / length)}' app.jsonl # Percentile approximation (p50, p95, p99) jq -sc ' map(.duration_ms) | sort | length as $n | { p50: .[($n * 0.50 | floor)], p95: .[($n * 0.95 | floor)], p99: .[($n * 0.99 | floor)], max: .[-1] } ' app.jsonl # Sum grouped by category jq -sc ' group_by(.service) | map({service: .[0].service, total_bytes: (map(.bytes) | add)}) ' app.jsonl ``` ### Group By with Aggregation ```bash # Group by service, show count and error rate jq -sc ' group_by(.service) | map({ service: .[0].service, total: length, errors: (map(select(.level == "error")) | length), error_rate: ((map(select(.level == "error")) | length) / length * 100 | round) }) ' app.jsonl # Group by hour jq -sc ' group_by(.timestamp | split("T")[1] | split(":")[0]) | map({ hour: .[0].timestamp | split("T")[1] | split(":")[0], count: length }) ' app.jsonl # Nested group by (service then level) jq -sc ' group_by(.service) | map({ service: .[0].service, by_level: (group_by(.level) | map({level: .[0].level, n: length})) }) ' app.jsonl ``` ### Top-N Queries ```bash # Top 10 slowest requests jq -sc 'sort_by(-.duration_ms) | .[:10] | .[] | {path: .path, ms: .duration_ms}' app.jsonl # Top 5 most frequent errors jq -sc ' map(select(.level == "error")) | group_by(.message) | map({message: .[0].message, count: length}) | sort_by(-.count) | .[:5] ' app.jsonl # Top users by request count jq -sc ' group_by(.user_id) | map({user: .[0].user_id, requests: length}) | sort_by(-.requests) | .[:10] ' app.jsonl ``` ### Histogram and Distribution Analysis ```bash # Response time histogram (buckets: 0-100, 100-500, 500-1000, 1000+) jq -sc ' map(.duration_ms) | { "0-100ms": (map(select(. < 100)) | length), "100-500ms": (map(select(. >= 100 and . < 500)) | length), "500-1000ms": (map(select(. >= 500 and . < 1000)) | length), "1000ms+": (map(select(. >= 1000)) | length) } ' app.jsonl # Status code distribution jq -rsc ' group_by(.status) | map("\(.[0].status)\t\(length)") | sort | .[] ' app.jsonl # Log level distribution over time (by hour) jq -rsc ' group_by(.timestamp | split("T")[1] | split(":")[0]) | map( (.[0].timestamp | split("T")[1] | split(":")[0]) as $hour | { hour: $hour, info: (map(select(.level == "info")) | length), warn: (map(select(.level == "warn")) | length), error: (map(select(.level == "error")) | length) } ) | .[] | [.hour, .info, .warn, .error] | @tsv ' app.jsonl | column -t -N HOUR,INFO,WARN,ERROR ``` ### Running Totals and Cumulative Sums ```bash # Cumulative error count over time jq -sc ' sort_by(.timestamp) | reduce .[] as $item ( {total: 0, rows: []}; .total += 1 | .rows += [{ts: $item.timestamp, cumulative: .total}] ) | .rows[] | [.ts, .cumulative] | @tsv ' <(jq -c 'select(.level == "error")' app.jsonl) # Running average of response times jq -sc ' sort_by(.timestamp) | foreach .[] as $item ( {n: 0, sum: 0}; .n += 1 | .sum += $item.duration_ms; {ts: $item.timestamp, running_avg: (.sum / .n | round)} ) ' app.jsonl ``` --- ## Transformation Patterns ### Reshape Objects ```bash # Flatten nested to flat jq -c '{ ts: .timestamp, level: .level, msg: .message, user: .context.user.id, method: .request.method, path: .request.path }' app.jsonl # Add computed fields jq -c '. + { date: (.timestamp | split("T")[0]), hour: (.timestamp | split("T")[1] | split(":")[0] | tonumber), is_error: (.level == "error") }' app.jsonl # Rename fields jq -c '{timestamp: .ts, message: .msg, severity: .lvl}' app.jsonl # Remove fields jq -c 'del(.stack_trace, .internal_debug_info)' app.jsonl ``` ### Merge Fields from Multiple Lines ```bash # Combine start and end events by request_id jq -sc ' group_by(.request_id) | map( (map(select(.event == "start")) | .[0]) as $start | (map(select(.event == "end")) | .[0]) as $end | { request_id: .[0].request_id, start: $start.timestamp, end: $end.timestamp, status: $end.status, path: $start.path } )[] ' events.jsonl # Merge consecutive lines (e.g., multiline log entries) jq -sc ' reduce .[] as $item ( []; if (. | length) == 0 then [$item] elif $item.continuation == true then (.[-1].message += "\n" + $item.message) | . else . + [$item] end )[] ' app.jsonl ``` ### Convert Between Formats ```bash # JSONL to CSV jq -r '[.timestamp, .level, .message] | @csv' app.jsonl > app.csv # JSONL to TSV jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl > app.tsv # JSONL to CSV with header echo "timestamp,level,message" > app.csv jq -r '[.timestamp, .level, .message] | @csv' app.jsonl >> app.csv # CSV to JSONL (using mlr) mlr --c2j cat app.csv > app.jsonl # JSONL to formatted table jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl | column -t -s$'\t' # JSONL to markdown table echo "| Timestamp | Level | Message |" echo "|-----------|-------|---------|" jq -r '"| \(.timestamp) | \(.level) | \(.message) |"' app.jsonl ``` ### Annotate Lines with Computed Fields ```bash # Add line number jq -c --argjson n 0 '. + {line_num: (input_line_number)}' app.jsonl # Add duration since previous event (requires slurp) jq -sc ' sort_by(.timestamp) | . as $all | [range(length)] | map( $all[.] + ( if . > 0 then {gap_from_prev: "computed"} else {gap_from_prev: null} end ) )[] ' app.jsonl # Tag lines matching criteria jq -c '. + { severity_class: ( if .level == "error" or .level == "fatal" then "critical" elif .level == "warn" then "warning" else "normal" end ) }' app.jsonl # Enrich with filename when processing multiple files fd -e jsonl . logs/ -x bash -c 'jq -c --arg src "$1" ". + {source: \$src}" "$1"' _ {} ``` --- ## Comparison Patterns ### Diff Two JSONL Files by Matching Key ```bash # Find entries in A but not in B (by id) jq -r '.id' b.jsonl | sort > /tmp/b_ids.txt jq -c --slurpfile bids <(jq -Rs 'split("\n") | map(select(. != ""))' /tmp/b_ids.txt) ' select(.id | IN($bids[0][])) | not ' a.jsonl # Simpler approach using comm jq -r '.id' a.jsonl | sort > /tmp/a_ids.txt jq -r '.id' b.jsonl | sort > /tmp/b_ids.txt comm -23 /tmp/a_ids.txt /tmp/b_ids.txt # IDs in A but not B comm -13 /tmp/a_ids.txt /tmp/b_ids.txt # IDs in B but not A comm -12 /tmp/a_ids.txt /tmp/b_ids.txt # IDs in both # Find records that exist in both but have different values jq -sc ' [., input] | (.[0] | map({(.id): .}) | add) as $a | (.[1] | map({(.id): .}) | add) as $b | ($a | keys) as $keys | [$keys[] | select($a[.] != $b[.])] | map({id: ., a: $a[.], b: $b[.]}) ' <(jq -sc '.' a.jsonl) <(jq -sc '.' b.jsonl) ``` ### Side-by-Side Field Comparison ```bash # Compare a specific field between two runs paste <(jq -r '[.id, .score] | @tsv' run1.jsonl | sort) \ <(jq -r '[.id, .score] | @tsv' run2.jsonl | sort) | awk -F'\t' '$2 != $4 {print $1, "run1=" $2, "run2=" $4}' # Summary comparison of two log files echo "=== File A ===" && jq -sc '{ lines: length, errors: (map(select(.level == "error")) | length), unique_users: ([.[].user_id] | unique | length) }' a.jsonl echo "=== File B ===" && jq -sc '{ lines: length, errors: (map(select(.level == "error")) | length), unique_users: ([.[].user_id] | unique | length) }' b.jsonl ``` ### Find New, Missing, and Changed Records ```bash # Comprehensive diff report jq -r '.id' a.jsonl | sort > /tmp/a.ids jq -r '.id' b.jsonl | sort > /tmp/b.ids echo "--- New in B (not in A) ---" comm -13 /tmp/a.ids /tmp/b.ids echo "--- Removed from A (not in B) ---" comm -23 /tmp/a.ids /tmp/b.ids echo "--- Changed (in both, different values) ---" comm -12 /tmp/a.ids /tmp/b.ids | while read id; do a_hash=$(rg "\"id\":\"$id\"" a.jsonl | md5sum | cut -d' ' -f1) b_hash=$(rg "\"id\":\"$id\"" b.jsonl | md5sum | cut -d' ' -f1) [ "$a_hash" != "$b_hash" ] && echo "$id" done ``` --- ## Performance Patterns ### Two-Stage rg + jq Pipeline The single most important performance pattern. ripgrep is 10-100x faster than jq at scanning text. ```bash # BAD: jq scans every line (slow on large files) jq -c 'select(.level == "error" and .service == "auth")' huge.jsonl # GOOD: rg filters text first, jq only parses matching lines rg '"error"' huge.jsonl | rg '"auth"' | jq -c '.' # GOOD: for precise matching after rg prefilter rg '"error"' huge.jsonl | jq -c 'select(.level == "error" and .service == "auth")' # Benchmarks (typical 1GB JSONL file): # jq alone: 45 seconds # rg + jq: 3 seconds # rg alone: 0.8 seconds ``` ### GNU parallel for Splitting Large Files ```bash # Split a 10GB file and process in parallel split -l 500000 huge.jsonl /tmp/chunk_ # Count errors across all chunks ls /tmp/chunk_* | parallel "rg -c '\"error\"' {}" | awk -F: '{sum+=$2} END {print sum}' # Extract and merge results ls /tmp/chunk_* | parallel "jq -c 'select(.level == \"error\")' {}" > all_errors.jsonl # Cleanup rm /tmp/chunk_* # One-liner with process substitution parallel --pipe -L 100000 'jq -c "select(.level == \"error\")"' < huge.jsonl > errors.jsonl ``` ### jq --stream for SAX-Style Processing For files too large to fit in memory, even line-by-line (e.g., a single 5GB JSON array). ```bash # Count items in a huge JSON array without loading it jq --stream 'select(.[0] | length == 1) | .[0][0]' huge-array.json | tail -1 # Extract specific field from each item in huge array jq -cn --stream 'fromstream(1 | truncate_stream(inputs)) | .name' huge-array.json # Filter items from huge array jq -cn --stream ' fromstream(1 | truncate_stream(inputs)) | select(.status == "failed") ' huge-array.json ``` ### Indexing Frequently-Queried Files ```bash # Build an index of line offsets by key value awk '{ match($0, /"request_id":"([^"]+)"/, m) if (m[1]) print m[1], NR }' app.jsonl | sort > app.idx # Look up specific request by index LINE=$(grep "req-abc-123" app.idx | awk '{print $2}') sed -n "${LINE}p" app.jsonl | jq . # Build a SQLite index for repeated queries sqlite3 log_index.db "CREATE TABLE idx (request_id TEXT, line INTEGER)" awk '{ match($0, /"request_id":"([^"]+)"/, m) if (m[1]) print "INSERT INTO idx VALUES (\047" m[1] "\047, " NR ");" }' app.jsonl | sqlite3 log_index.db # Query by index LINE=$(sqlite3 log_index.db "SELECT line FROM idx WHERE request_id = 'req-abc-123'") sed -n "${LINE}p" app.jsonl | jq . ``` ### Memory-Efficient Aggregation Without Slurp ```bash # Count by level without loading entire file jq -r '.level' app.jsonl | sort | uniq -c | sort -rn # Top error messages without slurp jq -r 'select(.level == "error") | .message' app.jsonl | sort | uniq -c | sort -rn | head -20 # Unique users without slurp jq -r '.user_id' app.jsonl | sort -u | wc -l # Sum without slurp jq -r '.bytes' app.jsonl | awk '{sum+=$1} END {print sum}' # These are all O(1) memory (streaming) vs O(n) memory (slurp) ``` -
tool-setup.md 17.9 KB
# Tool Setup Reference Installation, configuration, and key commands for log analysis tools. Each tool includes install commands for all platforms, the most useful flags, and integration patterns. --- ## jq -- JSON/JSONL Processor The primary tool for structured log analysis. Processes JSONL line by line (streaming) or as a batch (slurp). ### Installation ```bash # macOS brew install jq # Ubuntu/Debian sudo apt install jq # Windows choco install jq # or winget install jqlang.jq # Verify jq --version ``` ### Key Flags | Flag | Purpose | Example | |------|---------|---------| | `-c` | Compact output (one JSON per line) | `jq -c '.' file.jsonl` | | `-r` | Raw string output (no quotes) | `jq -r '.message' file.jsonl` | | `-s` | Slurp: read all lines into array | `jq -s 'length' file.jsonl` | | `-S` | Sort object keys | `jq -S '.' file.json` | | `-e` | Exit 1 if output is false/null | `jq -e '.ok' file.json` | | `-R` | Read lines as raw strings | `jq -R 'fromjson?' messy.jsonl` | | `-n` | Null input (use with inputs) | `jq -n '[inputs]' file.jsonl` | | `--arg` | Pass string variable | `jq --arg id "42" 'select(.id == $id)'` | | `--argjson` | Pass JSON variable | `jq --argjson n 10 'select(.count > $n)'` | | `--slurpfile` | Load file as variable | `jq --slurpfile ids ids.json 'select(.id | IN($ids[][]))'` | | `--stream` | SAX-style path/value output | `jq --stream '.' huge.json` | | `--unbuffered` | Flush after each output | `tail -f f.jsonl \| jq --unbuffered '.'` | | `--tab` | Use tabs for indentation | `jq --tab '.' file.json` | ### Essential Commands ```bash # Pretty print a single JSON object jq '.' file.json # Validate JSONL (report bad lines) jq -c '.' file.jsonl > /dev/null 2>&1 || echo "Invalid JSON detected" # Find and show invalid lines awk '{ cmd = "echo " "'\'''" $0 "'\'''" " | jq . 2>/dev/null" if (system(cmd) != 0) print NR": "$0 }' file.jsonl # Better: use jq -R to find invalid lines jq -R 'fromjson? // error' file.jsonl 2>&1 | rg "error" | head # Count lines in JSONL jq -sc 'length' file.jsonl # Get unique keys across all objects jq -sc '[.[] | keys[]] | unique' file.jsonl # Get schema (keys and types) from first line head -1 file.jsonl | jq '[to_entries[] | {key, type: (.value | type)}]' # Reformat JSONL with consistent key ordering jq -cS '.' file.jsonl > normalized.jsonl ``` ### Debugging jq Expressions ```bash # Use debug to print intermediate values to stderr jq '.items[] | debug | select(.active)' file.json # Use @text to see what jq thinks a value is jq '.field | @text' file.json # Use type to check value types jq '.field | type' file.json # Build expressions incrementally jq '.' file.json # Start: see full structure jq '.items' file.json # Navigate to array jq '.items[]' file.json # Iterate array jq '.items[] | .name' file.json # Extract field jq '.items[] | select(.active)' file.json # Filter # Common error: "Cannot iterate over null" # Fix: use ? operator jq '.items[]?' file.json # Won't error if items is null # Common error: "null is not iterable" # Fix: default empty array jq '(.items // [])[]' file.json ``` ### Integration with Other Tools ```bash # rg prefilter then jq parse rg '"error"' app.jsonl | jq -r '.message' # jq output to column for alignment jq -r '[.name, .status, .duration] | @tsv' app.jsonl | column -t # jq output to sort/uniq for frequency jq -r '.error_type' errors.jsonl | sort | uniq -c | sort -rn # jq to CSV for spreadsheet import jq -r '[.timestamp, .level, .message] | @csv' app.jsonl > export.csv # jq with xargs for per-line processing jq -r '.file_path' manifest.jsonl | xargs wc -l ``` --- ## lnav -- Log File Navigator Interactive terminal-based log viewer with SQL support, automatic format detection, timeline view, and filtering. Ideal for exploratory analysis. ### Installation ```bash # macOS brew install lnav # Ubuntu/Debian sudo apt install lnav # Windows (via Chocolatey) choco install lnav # From source curl -LO https://github.com/tstack/lnav/releases/download/v0.12.2/lnav-0.12.2-linux-musl-x86_64.zip unzip lnav-0.12.2-linux-musl-x86_64.zip sudo cp lnav-0.12.2/lnav /usr/local/bin/ # Verify lnav -V ``` ### Key Features | Feature | Access | Description | |---------|--------|-------------| | Auto-detect format | Automatic | Recognizes syslog, Apache, nginx, JSON, and many more | | SQL queries | `:` then SQL | Run SQL against log data | | Filter in/out | `i` / `o` | Interactive include/exclude filters | | Bookmarks | `m` | Mark lines for later reference | | Timeline | `t` | Show time histogram | | Pretty print | `p` | Toggle pretty-printing JSON | | Headless mode | `-n -c "..."` | Non-interactive command execution | | Compressed files | Automatic | Handles .gz, .bz2, .xz transparently | ### Essential Commands ```bash # Open log file(s) lnav app.log lnav /var/log/syslog /var/log/auth.log # multiple files, merged by timestamp # Open JSONL logs lnav app.jsonl # Open compressed logs lnav app.log.gz # Open all logs in a directory lnav /var/log/myapp/ # Headless mode: run query and output results lnav -n -c ";SELECT count(*) FROM logline WHERE log_level = 'error'" app.log # Headless mode: filter and export lnav -n -c ";SELECT log_time, log_body FROM logline WHERE log_level = 'error'" \ -c ":write-csv-to errors.csv" app.log # Headless mode: get stats lnav -n -c ";SELECT log_level, count(*) as cnt FROM logline GROUP BY log_level ORDER BY cnt DESC" app.log ``` ### SQL Mode Recipes ```sql -- Error count by hour SELECT strftime('%Y-%m-%d %H', log_time) as hour, count(*) as errors FROM logline WHERE log_level = 'error' GROUP BY hour ORDER BY hour; -- Top error messages SELECT log_body, count(*) as cnt FROM logline WHERE log_level = 'error' GROUP BY log_body ORDER BY cnt DESC LIMIT 10; -- Time between events SELECT log_time, log_body, julianday(log_time) - julianday(lag(log_time) OVER (ORDER BY log_time)) as gap_days FROM logline WHERE log_level = 'error'; -- Log volume over time SELECT strftime('%Y-%m-%d %H:%M', log_time) as minute, count(*) as lines FROM logline GROUP BY minute ORDER BY minute; ``` ### Custom Log Formats ```json // ~/.lnav/formats/installed/myapp.json { "myapp_log": { "title": "My Application Log", "regex": { "std": { "pattern": "^(?<timestamp>\\d{4}-\\d{2}-\\d{2}T\\d{2}:\\d{2}:\\d{2})\\s+\\[(?<level>\\w+)\\]\\s+(?<body>.*)" } }, "timestamp-format": ["%Y-%m-%dT%H:%M:%S"], "level": { "error": "ERROR", "warning": "WARN", "info": "INFO", "debug": "DEBUG" } } } ``` ### Interactive Keyboard Shortcuts | Key | Action | |-----|--------| | `/` | Search forward (regex) | | `n` / `N` | Next/previous search match | | `i` | Toggle filter: show only matching lines | | `o` | Toggle filter: hide matching lines | | `TAB` | Switch between views (log, text, help) | | `t` | Toggle timeline histogram | | `m` | Set bookmark on current line | | `u` / `U` | Next/previous bookmark | | `z` / `Z` | Zoom in/out on timeline | | `p` | Toggle pretty-print for JSON | | `e` / `E` | Next/previous error | | `w` / `W` | Next/previous warning | | `:` | Enter command mode | | `;` | Enter SQL query mode | --- ## angle-grinder (agrind) -- Log Pipeline Aggregation Pipeline-based aggregation tool designed for log analysis. Think SQL-like queries in a streaming pipeline syntax. ### Installation ```bash # Via cargo (all platforms) cargo install ag # macOS brew install angle-grinder # Verify agrind --version ``` ### Pipeline Syntax ``` <input_pattern> | <operator1> | <operator2> | ... ``` ### Essential Commands ```bash # Count log levels cat app.log | agrind '* | parse "* [*] *" as ts, level, msg | count by level' # Top URLs cat access.log | agrind '* | parse "* * * * * * *" as ip, _, _, ts, method, url, status | count by url | sort by _count desc | head 10' # Average response time by endpoint cat access.log | agrind '* | parse "* *ms" as prefix, duration | avg of duration by prefix' # Error frequency over time cat app.log | agrind '* | parse "*T*:*:* [ERROR]*" as date, hour, min, sec, msg | count by hour' # Filter then aggregate cat app.log | agrind '* | where level == "error" | count by msg | sort by _count desc' # JSON log fields cat app.jsonl | agrind '* | json | where level == "error" | count by message' ``` ### Operators Reference | Operator | Purpose | Example | |----------|---------|---------| | `parse` | Extract fields with pattern | `parse "* [*] *" as a, b, c` | | `json` | Parse JSON log lines | `json` | | `where` | Filter rows | `where level == "error"` | | `count` | Count (optionally by group) | `count by level` | | `sum` | Sum a field | `sum of bytes` | | `avg` | Average a field | `avg of duration` | | `min` / `max` | Min/max of field | `min of response_time` | | `sort` | Sort results | `sort by _count desc` | | `head` | Limit results | `head 10` | | `uniq` | Unique values | `uniq by user_id` | | `percentile` | Percentile calc | `p50 of duration, p99 of duration` | --- ## rg (ripgrep) -- Fast Pattern Search Already covered extensively in file-search skill. Here are log-specific flags and patterns. ### Log-Specific Flags | Flag | Purpose | Example | |------|---------|---------| | `-c` | Count matches per file | `rg -c "ERROR" /var/log/*.log` | | `-l` | List files with matches | `rg -l "timeout" /var/log/` | | `-L` | List files without matches | `rg -L "healthy" /var/log/` | | `--stats` | Show match statistics | `rg --stats "error" app.log` | | `-A N` | Show N lines after match | `rg -A5 "Exception" app.log` | | `-B N` | Show N lines before match | `rg -B3 "FATAL" app.log` | | `-C N` | Show N lines context | `rg -C5 "crash" app.log` | | `-U` | Multiline matching | `rg -U "Error.*\n.*at " app.log` | | `--json` | JSON output format | `rg --json "error" app.log` | | `-a` | Search binary files | `rg -a "pattern" binary.log` | | `--line-buffered` | Flush per line (for tail) | `tail -f app.log \| rg --line-buffered "error"` | | `-F` | Fixed string (no regex) | `rg -F "stack[0]" app.log` | | `-f FILE` | Patterns from file | `rg -f patterns.txt app.log` | | `-v` | Invert match | `rg -v "DEBUG" app.log` | ### Log Search Recipes ```bash # Find errors across all log files recursively rg "ERROR|FATAL|CRITICAL" /var/log/ # Count errors per file, sorted rg -c "ERROR" /var/log/ 2>/dev/null | sort -t: -k2 -rn # Find stack traces (multiline) rg -U "Exception.*\n(\s+at .*\n)+" app.log # Extract timestamps of errors rg "ERROR" app.log | rg -o "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}" # Search compressed log files rg -z "error" app.log.gz # Search JSONL for specific field value (text-level, fast but approximate) rg '"level":"error"' app.jsonl # Search JSONL for value in specific key (avoid matching wrong key) rg '"user_id":"user-42"' app.jsonl # Negative lookahead: errors that are NOT timeouts rg "ERROR(?!.*timeout)" app.log # Time-bounded search (extract lines between two timestamps) rg "2026-03-08T1[4-5]:" app.log ``` ### rg JSON Output Mode ```bash # Get structured output from rg (useful for programmatic processing) rg --json "error" app.log | jq -c 'select(.type == "match") | {file: .data.path.text, line: .data.line_number, text: .data.lines.text}' # Count matches with file info rg --json "error" app.log | jq -c 'select(.type == "summary") | .data.stats' ``` --- ## awk -- Column-Based Log Processing Pre-installed on all Unix systems. Best for space/tab delimited logs with consistent column structure. ### Common Recipes ```bash # Apache/nginx combined log format columns: # $1=IP $2=ident $3=user $4=date $5=time $6=tz $7=method $8=path $9=proto $10=status $11=size # Status code distribution awk '{print $9}' access.log | sort | uniq -c | sort -rn # Requests per IP awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -20 # 5xx errors with paths awk '$9 >= 500 {print $1, $9, $7}' access.log # Average response size awk '{sum += $10; n++} END {printf "Avg: %.0f bytes\n", sum/n}' access.log # Requests per minute awk '{print substr($4, 2, 17)}' access.log | sort | uniq -c | tail -20 # Bandwidth by path awk '{bytes[$7] += $10} END {for (p in bytes) printf "%10d %s\n", bytes[p], p}' access.log | sort -rn | head -20 # Custom delimiter (e.g., pipe-separated) awk -F'|' '{print $3, $5}' custom.log # Time-range filter (syslog format) awk '/^Mar 8 14:/ {print}' syslog # Calculate time difference between first and last line awk 'NR==1 {first=$1" "$2} END {last=$1" "$2; print "From:", first, "To:", last}' app.log ``` ### awk for Key-Value Logs ```bash # Parse key=value format (logfmt) awk '{ for (i=1; i<=NF; i++) { split($i, kv, "=") if (kv[1] == "duration") sum += kv[2]; n++ } } END {print "avg_duration=" sum/n}' app.log # Extract specific key from logfmt awk '{ for (i=1; i<=NF; i++) { split($i, kv, "=") if (kv[1] == "status" && kv[2] >= 500) print $0 } }' app.log ``` --- ## GNU parallel -- Parallel Log Processing Splits work across CPU cores for processing large log files. ### Installation ```bash # macOS brew install parallel # Ubuntu/Debian sudo apt install parallel # Verify parallel --version ``` ### Essential Commands ```bash # Process multiple log files in parallel ls /var/log/app-*.jsonl | parallel "jq -c 'select(.level == \"error\")' {} > {.}_errors.jsonl" # Split large file and process chunks in parallel split -l 200000 huge.jsonl /tmp/chunk_ ls /tmp/chunk_* | parallel "jq -r '.message' {} | sort | uniq -c" | sort -rn | head -20 # Parallel grep across many files fd -e jsonl . /var/log/ | parallel "rg -c '\"error\"' {} 2>/dev/null" | sort -t: -k2 -rn # Pipe-based parallelism (no temp files) parallel --pipe -L 50000 "jq -c 'select(.level == \"error\")'" < huge.jsonl > errors.jsonl # Parallel with progress bar ls /var/log/app-*.jsonl | parallel --bar "jq -sc 'length' {}" | awk '{sum+=$1} END {print sum, "total lines"}' # Number of jobs (default: CPU cores) ls *.jsonl | parallel -j 4 "jq -c 'select(.status >= 500)' {}" ``` ### Combining with split ```bash # Full workflow: split, process in parallel, merge results FILE=huge.jsonl CHUNKS=/tmp/log_chunks mkdir -p "$CHUNKS" # Split split -l 100000 "$FILE" "$CHUNKS/chunk_" # Process in parallel ls "$CHUNKS"/chunk_* | parallel "jq -r 'select(.level == \"error\") | .message' {}" | sort | uniq -c | sort -rn > error_summary.txt # Cleanup rm -rf "$CHUNKS" ``` --- ## Miller (mlr) -- CSV/TSV Log Analysis Like awk, sed, and jq combined but specifically for structured record data (CSV, TSV, JSON). ### Installation ```bash # macOS brew install miller # Ubuntu/Debian sudo apt install miller # Windows choco install miller # Verify mlr --version ``` ### Essential Commands ```bash # View CSV with headers mlr --csv head -n 10 access_log.csv # Filter rows mlr --csv filter '$status >= 400' access_log.csv # Sort by column mlr --csv sort-by -nr duration access_log.csv # Statistics mlr --csv stats1 -a min,max,mean,p95 -f duration access_log.csv # Group by with stats mlr --csv stats1 -a count,mean -f duration -g endpoint access_log.csv # Convert formats mlr --c2j cat access_log.csv # CSV to JSON mlr --c2t cat access_log.csv # CSV to TSV (table) mlr --j2c cat access_log.json # JSON to CSV mlr --c2p cat access_log.csv # CSV to pretty-print table # Top-N by group mlr --csv top -n 5 -f duration -g endpoint access_log.csv # Add computed fields mlr --csv put '$error = ($status >= 400 ? "yes" : "no")' access_log.csv # Decimate (sample every Nth row) mlr --csv sample -k 100 huge_log.csv # Uniq count mlr --csv count-distinct -f status access_log.csv # Histogram mlr --csv decimate -g status -n 1 access_log.csv | mlr --csv count-distinct -f status # Join two CSV files mlr --csv join -j user_id -f users.csv then sort-by user_id access_log.csv ``` ### TSV from jq to mlr Pipeline ```bash # Extract JSONL to TSV, then use mlr for analysis jq -r '[.timestamp, .level, .duration_ms, .path] | @tsv' app.jsonl > /tmp/extracted.tsv mlr --tsvlite --from /tmp/extracted.tsv \ label timestamp,level,duration_ms,path then \ filter '$level == "error"' then \ stats1 -a count,mean -f duration_ms -g path then \ sort-by -nr count ``` --- ## Tool Integration Cheat Sheet ### Combining Tools ```bash # fd + rg + jq: find files, prefilter, extract fd -e jsonl . logs/ | xargs rg -l '"error"' | xargs jq -c 'select(.level == "error") | {ts: .timestamp, msg: .message}' # rg + jq + column: search, extract, format rg '"timeout"' app.jsonl | jq -r '[.timestamp, .service, .message] | @tsv' | column -t # jq + sort + uniq: aggregate without slurp jq -r '.error_type' errors.jsonl | sort | uniq -c | sort -rn # tail + rg + jq: live monitoring with extraction tail -f app.jsonl | rg --line-buffered '"error"' | jq --unbuffered -r '[.timestamp, .message] | @tsv' # fd + parallel + jq: parallel extraction across many files fd -e jsonl . logs/ | parallel "jq -c 'select(.level == \"error\")' {}" > all_errors.jsonl # jq + mlr: structured extraction then statistical analysis jq -r '[.path, .duration_ms, .status] | @csv' app.jsonl | \ mlr --csv label path,duration,status then \ stats1 -a p50,p95,p99 -f duration -g path then \ sort-by -nr p95 # lnav + headless SQL: non-interactive queries lnav -n -c ";SELECT log_level, count(*) FROM logline GROUP BY log_level" app.log ``` ### Decision Guide: Which Combination? ``` Task: Explore unknown log file --> lnav (interactive, auto-detects format) Task: Quick search for pattern --> rg "pattern" file.log Task: Extract fields from JSONL --> jq -r '[.field1, .field2] | @tsv' file.jsonl Task: Count/aggregate JSONL (<100MB) --> jq -sc 'group_by(.x) | map(...)' file.jsonl Task: Count/aggregate JSONL (>100MB) --> jq -r '.field' file.jsonl | sort | uniq -c | sort -rn Task: Search large JSONL then extract --> rg "pattern" file.jsonl | jq -r '.field' Task: CSV/TSV log statistics --> mlr --csv stats1 -a mean,p95 -f duration file.csv Task: Process many log files in parallel --> fd -e jsonl . dir/ | parallel "jq ..." Task: Pipeline aggregation on text logs --> cat file.log | agrind '* | parse ... | count by ...' Task: Live monitoring with filtering --> tail -f file.jsonl | rg --line-buffered "x" | jq --unbuffered '.' ```
-
-
scripts
-
.gitkeep 0 B · in bundle
-
-
SKILL.md 15.8 KB
--- name: log-ops description: "Log analysis and JSONL processing - structured extraction, cross-log correlation, timeline reconstruction, pattern search" license: MIT allowed-tools: "Read Edit Write Bash Glob Grep Agent" metadata: author: claude-mods related-skills: data-processing, debug-ops, monitoring-ops, file-search, introspect --- # Log Operations Practical patterns for analyzing log files -- especially JSONL format used in agent conversation logs, benchmark outputs, and structured application logs. ## Log Format Decision Tree ``` Unknown Log File │ ├─ Is it one JSON object per line? │ ├─ Yes ──────────────────────── JSONL │ │ ├─ Small file (<100MB) │ │ │ └─ jq for extraction, jq -s for aggregation │ │ ├─ Large file (100MB-1GB) │ │ │ └─ rg prefilter then pipe to jq │ │ └─ Huge file (>1GB) │ │ └─ split + parallel jq, or jq --stream │ │ │ └─ No │ ├─ Is it one large JSON object/array? │ │ └─ Yes ──────────────── Single JSON │ │ └─ jq --stream for SAX-style, or jq directly if fits in memory │ │ │ ├─ Does it have key=value pairs? │ │ └─ Yes ──────────────── Structured (logfmt / key-value) │ │ └─ rg for search, awk/sd for extraction, angle-grinder for aggregation │ │ │ ├─ Does it follow syslog format? (timestamp hostname service[pid]: message) │ │ └─ Yes ──────────────── Syslog │ │ └─ rg for search, awk for column extraction, lnav for interactive │ │ │ ├─ Is it space/tab delimited with consistent columns? │ │ └─ Yes ──────────────── Column-based (access logs, CSV) │ │ └─ awk for extraction, mlr for CSV, rg for pattern search │ │ │ └─ Mixed or unstructured │ └─ Plain text ─────────── Freeform │ └─ rg for search, rg -A/-B for context, lnav for exploration ``` ## Prerequisites **Required** (must be installed): - `rg` (ripgrep) - text search, prefiltering. Install: `cargo install ripgrep` / `choco install ripgrep` - `jq` - JSON/JSONL extraction and transformation. Install: `brew install jq` / `choco install jq` **Optional** (enhanced capabilities, gracefully degraded without): - `lnav` - interactive log exploration with SQL queries. Install: `brew install lnav` / WSL: `apt install lnav` - `agrind` (angle-grinder) - pipeline aggregation syntax. Install: `cargo install ag` - `mlr` (Miller) - CSV/TSV log analysis. Install: `brew install miller` / `choco install miller` - `GNU parallel` - parallel processing of split files. Install: `brew install parallel` > All patterns in this skill work with just rg + jq. Optional tools add interactive exploration (lnav), pipeline aggregation (agrind), and tabular analysis (mlr). ## Tool Selection Matrix | Tool | Best For | Speed | Required? | |------|----------|-------|-----------| | `rg` (ripgrep) | Raw pattern matching in any format | Fastest | Yes | | `jq` | JSONL structured extraction and transformation | Fast | Yes | | `jq -s` | JSONL aggregation (slurp all lines into array) | Medium (loads all into memory) | Yes (part of jq) | | `lnav` | Interactive exploration, SQL over logs | Interactive | Optional | | `agrind` (angle-grinder) | Pipeline aggregation and counting | Fast | Optional | | `awk` | Column-based log formats, field extraction | Fast | Pre-installed | | `mlr` (Miller) | CSV/TSV log analysis, statistics | Fast | Optional | | `fd` + `rg` | Searching across many log directories | Fast | Pre-installed in dev-shell | | `GNU parallel` | Splitting large files for parallel processing | N/A (orchestrator) | Optional | ### When to Use What ``` Need to... │ ├─ Find lines matching a pattern │ └─ rg (always fastest for text search) │ ├─ Extract specific fields from JSONL │ └─ jq -r '[.field1, .field2] | @tsv' │ ├─ Count/aggregate over JSONL │ └─ jq -sc 'group_by(.field) | map({key: .[0].field, n: length})' │ ├─ Search JSONL by value then format results │ └─ rg '"error"' file.jsonl | jq -r '.message' (two-stage) │ ├─ Explore interactively with filtering/SQL │ └─ lnav file.log │ ├─ Aggregate with pipeline syntax │ └─ agrind '* | parse "* * *" as ts, level, msg | count by level' │ ├─ Extract columns from space-delimited logs │ └─ awk '{print $1, $4, $7}' access.log │ └─ Process CSV/TSV logs with headers └─ mlr --csv filter '$status >= 400' then stats1 -a count -f status ``` ## JSONL Quick Reference The most common format for structured logs. One JSON object per line, no trailing commas, no wrapping array. ### Stream Filtering (line by line, constant memory) ```bash # Filter by field value jq -c 'select(.level == "error")' app.jsonl # Filter by nested field jq -c 'select(.request.method == "POST")' app.jsonl # Filter by multiple conditions jq -c 'select(.level == "error" and .status >= 500)' app.jsonl # Filter by array contains jq -c 'select(.tags | index("critical"))' app.jsonl # Filter by field existence jq -c 'select(.stack_trace != null)' app.jsonl # Negate a filter jq -c 'select(.level != "debug")' app.jsonl ``` ### Field Extraction ```bash # Extract single field jq -r '.message' app.jsonl # Extract multiple fields as TSV jq -r '[.timestamp, .level, .message] | @tsv' app.jsonl # Extract with default for missing fields jq -r '.error_code // "none"' app.jsonl # Extract nested field safely jq -r '.response.headers["content-type"] // "unknown"' app.jsonl ``` ### Aggregation (requires slurp: loads entire file) ```bash # Count by field value jq -sc 'group_by(.level) | map({level: .[0].level, count: length})' app.jsonl # Top-N most common values jq -sc '[.[].error_type] | group_by(.) | map({type: .[0], count: length}) | sort_by(-.count) | .[:10]' app.jsonl # Sum a numeric field jq -sc 'map(.duration_ms) | add' app.jsonl # Average jq -sc 'map(.duration_ms) | add / length' app.jsonl # Min and max jq -sc 'map(.duration_ms) | {min: min, max: max}' app.jsonl ``` ### Nested Extraction (agent logs, complex structures) ```bash # Extract tool calls from conversation logs jq -c '.content[]? | select(.type == "tool_use") | .name' conversation.jsonl # De-escape nested JSON strings jq -c '.content | fromjson' app.jsonl # Flatten nested arrays jq -c '[.events[]? | .action]' app.jsonl # Extract from arrays of objects jq -c '.results[]? | select(.passed == false) | {test: .name, error: .message}' results.jsonl ``` ### Two-Stage Pipeline (rg for speed, jq for structure) ```bash # Fast prefilter then structured extraction rg '"error"' app.jsonl | jq -r '[.timestamp, .message] | @tsv' # Search for specific value then aggregate rg '"timeout"' app.jsonl | jq -sc 'length' # Pattern match then extract rg '"user_id":"u-123"' app.jsonl | jq -c '{ts: .timestamp, action: .action}' ``` ### Time-Range Filtering ```bash # Filter by timestamp range (ISO 8601 string comparison works) jq -c 'select(.timestamp > "2026-03-08T10:00" and .timestamp < "2026-03-08T11:00")' app.jsonl # Events in the last N minutes (using epoch seconds) jq -c --arg cutoff "$(date -d '30 minutes ago' +%s)" 'select((.timestamp | sub("\\.[0-9]+Z$"; "Z") | fromdate) > ($cutoff | tonumber))' app.jsonl # Extract hour for histogram jq -r '.timestamp | split("T")[1] | split(":")[0]' app.jsonl | sort | uniq -c ``` ### Cross-File Join ```bash # Extract IDs from one file, search in another jq -r '.request_id' errors.jsonl | while read id; do rg "\"$id\"" responses.jsonl | jq -c '{id: .request_id, status: .status}' done # Faster: build lookup, then join jq -r '.request_id' errors.jsonl | sort -u > /tmp/error_ids.txt rg -Ff /tmp/error_ids.txt responses.jsonl | jq -c '{id: .request_id, status: .status}' # Join two JSONL files by key using jq --slurpfile jq --slurpfile lookup <(jq -sc 'map({(.id): .}) | add' lookup.jsonl) \ '. + ($lookup[0][.ref_id] // {})' main.jsonl ``` ## Plain Text Log Patterns ### Pattern Search with Context ```bash # Show 5 lines before and after each match rg -B5 -A5 "OutOfMemoryError" app.log # Show only matching files rg -l "FATAL" /var/log/ # Count matches per file rg -c "ERROR" /var/log/*.log | sort -t: -k2 -rn # Multiline patterns (stack traces) rg -U "Exception.*\n(\s+at .*\n)+" app.log ``` ### Column Extraction with awk ```bash # Apache/nginx access log: extract status codes awk '{print $9}' access.log | sort | uniq -c | sort -rn # Extract specific time range from syslog awk '$0 >= "Mar 8 10:00" && $0 <= "Mar 8 11:00"' syslog # Calculate average response time (column 11) awk '{sum += $11; n++} END {print sum/n}' access.log # Filter by status code and show URL + response time awk '$9 >= 500 {print $7, $11"ms"}' access.log ``` ### Live Monitoring ```bash # Follow with filtering tail -f app.log | rg --line-buffered "ERROR" # Follow JSONL and extract fields tail -f app.jsonl | jq --unbuffered -r '[.timestamp, .level, .message] | @tsv' # Follow multiple files tail -f /var/log/service-*.log | rg --line-buffered "error|warn" ``` ## Timeline Reconstruction ### Extracting and Sorting by Timestamp ```bash # Merge multiple log files by timestamp sort -t' ' -k1,2 service-a.log service-b.log > timeline.log # JSONL: sort by timestamp field jq -sc 'sort_by(.timestamp)[]' combined.jsonl > sorted.jsonl # Extract timestamps and calculate gaps jq -r '.timestamp' app.jsonl | awk ' NR > 1 { cmd = "date -d \"" prev "\" +%s"; cmd | getline t1; close(cmd) cmd = "date -d \"" $0 "\" +%s"; cmd | getline t2; close(cmd) gap = t2 - t1 if (gap > 5) print gap "s gap before " $0 } { prev = $0 } ' # Quick duration between first and last event jq -sc '{start: .[0].timestamp, end: .[-1].timestamp}' app.jsonl ``` ### Calculating Durations Between Events ```bash # Duration between paired events (start/end) jq -sc ' group_by(.request_id) | map( (map(select(.event == "start")) | .[0].timestamp) as $start | (map(select(.event == "end")) | .[0].timestamp) as $end | {id: .[0].request_id, start: $start, end: $end} ) ' events.jsonl # Identify the slowest phase jq -sc ' sort_by(.timestamp) | [range(1; length) | { from: .[.-1].event, to: .[.].event, gap: ((.[.].ts_epoch) - (.[.-1].ts_epoch)) }] | sort_by(-.gap) | .[0] ' events.jsonl ``` ## Cross-Log Correlation ### By Correlation ID ```bash # Find a request across all service logs fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {} # Build a timeline for a single request fd -e jsonl . /var/log/services/ -x rg "\"req-abc-123\"" {} \; | jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .event] | @tsv' ``` ### By Timestamp Window ```bash # Find events within 2 seconds of a known event # First get the target timestamp TARGET="2026-03-08T14:23:15" jq -c --arg t "$TARGET" ' select( .timestamp > ($t | sub("15$"; "13")) and .timestamp < ($t | sub("15$"; "17")) ) ' other-service.jsonl ``` ### By Session/User ```bash # Reconstruct a user session across log files fd -e jsonl . /var/log/ -x rg "\"user-42\"" {} \; | jq -sc 'sort_by(.timestamp)[] | [.timestamp, .service, .action] | @tsv' ``` ## Large File Strategies ### Search Recent Only ```bash # Last 10,000 lines (fast for append-only logs) tail -n 10000 huge.log | rg "pattern" # Last N lines of JSONL with structured extraction tail -n 5000 huge.jsonl | jq -c 'select(.level == "error")' ``` ### Split for Parallel Processing ```bash # Split into 100K-line chunks split -l 100000 huge.jsonl /tmp/chunk_ # Process in parallel fd 'chunk_' /tmp/ -x jq -c 'select(.level == "error")' {} > errors.jsonl # With GNU parallel split -l 100000 huge.jsonl /tmp/chunk_ ls /tmp/chunk_* | parallel 'jq -c "select(.level == \"error\")" {} >> /tmp/errors.jsonl' ``` ### Streaming for Huge Single JSON ```bash # SAX-style processing of a huge JSON array jq --stream 'select(.[0][0] == "results" and .[0][-1] == "status") | .[1]' huge.json # Extract items from a huge array without loading all jq -cn --stream 'fromstream(1 | truncate_stream(inputs))' huge-array.json ``` ### Two-Stage Always ```bash # ALWAYS faster: rg filters text, jq parses survivors rg '"error"' huge.jsonl | jq -r '.message' # vs. SLOW: jq reads and parses every line jq -r 'select(.level == "error") | .message' huge.jsonl ``` ## Search Across Directories ### Multi-Directory Patterns ```bash # Find all JSONL files with errors across trial directories fd -e jsonl . trials/ -x rg -l '"error"' {} # Count errors per log file across directories fd -e jsonl . trials/ -x bash -c 'echo "$(rg -c "\"error\"" "$1" 2>/dev/null || echo 0) $1"' _ {} # Extract and aggregate across directories fd -e jsonl . trials/ -x jq -c 'select(.level == "error") | {file: input_filename, msg: .message}' {} # Build summary table from multiple runs for dir in trials/*/; do total=$(wc -l < "$dir/results.jsonl") errors=$(rg -c '"error"' "$dir/results.jsonl" 2>/dev/null || echo 0) echo -e "$dir\t$total\t$errors" done | column -t -N DIR,TOTAL,ERRORS ``` ## Common Gotchas | Gotcha | Why It Hurts | Fix | |--------|-------------|-----| | `jq -s` on huge files loads everything into memory | OOM crash or swap thrashing on files over ~500MB | Use streaming: `rg` prefilter, `jq --stream`, or `split` + parallel | | JSONL with embedded newlines in string values | Line-by-line tools (rg, awk, head) split a single record across lines | Use `jq -c` to re-compact, or `jq -R 'fromjson?'` to skip malformed lines | | rg matches JSON keys, not just values | `rg "error"` matches `{"error_count": 0}` which is not an error | Use `rg '"level":"error"'` or pipe to `jq 'select(.level == "error")'` | | Timezone mismatches in timestamp comparisons | Events appear out of order or time ranges miss data | Normalize to UTC before comparing: `jq '.timestamp |= sub("\\+.*"; "Z")'` | | Unicode and escape sequences in log messages | jq chokes on invalid UTF-8 or double-escaped strings | Prefilter with `rg -a` (binary mode), or use `jq -R` for raw strings | | Inconsistent JSON schemas across log lines | `jq` errors on lines missing expected fields | Use `//` operator for defaults: `.field // "missing"` and `?` for optional: `.arr[]?` | | Forgetting `-c` flag with jq on JSONL | jq pretty-prints each line, output is no longer valid JSONL | Always use `jq -c` when output feeds into another JSONL consumer | | tail -f with jq buffering | Output appears delayed or not at all | Use `jq --unbuffered` or `stdbuf -oL jq` | | Sorting JSONL by timestamp without slurp | `sort` command does lexicographic sort on whole lines, not by field | Either `jq -sc 'sort_by(.timestamp)[]'` or extract timestamp prefix first | | Assuming log files are complete | Logs may be rotated, compressed, or still being written | Check for `.gz` rotated files: `fd -e gz . /var/log/ -x zcat {} \| rg pattern` | | Single quotes in jq on Windows | PowerShell/cmd do not handle single quotes the same as bash | Use double quotes with escaped inner quotes, or write jq filter to a file | ## Reference Files | File | Contents | Lines | |------|----------|-------| | `references/jsonl-patterns.md` | JSONL extraction, aggregation, transformation, comparison, and performance patterns | ~700 | | `references/analysis-workflows.md` | Agent conversation analysis, application log analysis, benchmark result parsing, cross-directory workflows | ~600 | | `references/tool-setup.md` | Installation and configuration for jq, lnav, angle-grinder, rg, awk, GNU parallel, Miller | ~450 | ## See Also - **data-processing** -- JSON/YAML/TOML processing with jq and yq - **debug-ops** -- Systematic debugging methodology, log-based debugging section - **monitoring-ops** -- Production observability, alerting, dashboards - **file-search** -- Finding files with fd, searching code with rg
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.