debug-methodology
Use when debugging embedded firmware issues — analyzing crash logs, investigating state machine lockups, tracing sensor/signal anomalies, or performing structured root-cause analysis after a crash has been located. For fault-register and stack-frame triage, load hardfault-triage
Install
npx skills add https://github.com/AmethystLuna/embedded-workbench/tree/master/skills/debug-methodology
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install amethystluna-embedded-workbench@llmmart
git clone https://github.com/AmethystLuna/embedded-workbench.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole amethystluna/embedded-workbench collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Debug Methodology
Debug by tracing values, not symptoms. These patterns come from real debugging sessions where surface-level fixes failed and root-cause analysis succeeded.
REQUIRED BACKGROUND: If the issue involves state machines or protocol timeouts, load Skill("state-machine-design") first. If the issue involves FreeRTOS tasks, ISRs, or NVM storage, load Skill("embedded-firmware-dev") first. Understand the domain rules before applying debugging methodology.
Red Flags
| You think | Reality |
|---|---|
| "I know where the bug is, let me fix it" | You know the symptom location. The root cause is often 3 layers away in a different module. |
| "One more round of trial fixes and I'll get it" | After 2 failed attempts, you need methodology, not persistence. |
| "I'll just add a bounds check and call it done" | You're masking a symptom. The real fix is at the data source, not at every consumer. |
| "The logs look normal, it must be a hardware glitch" | If you haven't correlated timestamps to code branches, you haven't actually read the logs. |
Iron Rules
Log first, not code first: Correlate serial log timestamps to code branches before touching code. One log line at a known timestamp is worth more than reading five source files.
Call-point census: When a value isn't updating, find all call sites of the update function. A single call site (e.g.,
sensor_update()only atmain_loop.c:47) immediately explains why refresh is delayed or gated by irrelevant conditions.grep -rn "sensor_update" --include="*.c" src/main_loop.c:128: sensor_update(&g_pressure);Only two call sites. If
main_loop.cis gated onwifi_connected, the sensor won't refresh until WiFi connects — that's the root cause.Cache freshness ≠ source data readiness: In embedded systems, "sensor has data" and "derived cache is refreshed" are concurrent events. Prefer pull mode (sync cache on read when source is ready). Avoid pure push mode (periodic background updaters may not fire before the first consumer reads).
// BAD: push mode — timer callback pushes data before consumer asks for it static void sensor_timer_cb(TimerHandle_t xTimer) { g_sensor.cache = sensor_read_raw(); // Timer owns refresh timing } float get_temperature(void) { return g_sensor.cache.temperature; // Stale if timer hasn't fired yet } // GOOD: pull mode — consumer triggers refresh when source is ready float get_temperature(void) { if (sensor_is_ready()) { sensor_sync_cache(&g_sensor); // Refresh on read } return g_sensor.cache.temperature; // As fresh as hardware allows }Multi-path convergence: When multiple independent code paths show the same error, find the shared state or cache they all read. Fix once at the update point — smaller fix, and future callers can't bypass it.
Ownership boundary mapping: Before changing behavior, identify which module owns the truth, which derives policy, and which only consumes. Don't let the GUI control backlight policy. Don't let the backlight module control DND state.
Progressive narrowing: Each investigation round shrinks scope — phenomenon → mechanism → specific state → root cause. Don't try to solve everything at once.
Minimal root-cause fix: The fix is usually 1-2 lines at the data source. If you're changing 5+ call sites, stop and ask: what single state change would make all of them correct without modification?
Library source is truth: After 2-3 rounds of custom implementation failure, stop iterating. Read the library source code (e.g., LVGL's
lv_line.c,lv_chart.c) to understand the native mechanism. Adopt and adapt. Verified patterns beat custom math.
Fix Principles
- Draw the failure signature, event timeline, and state transition chain before changing code.
- For defects near state transitions: trace the exact state consumed by that output. Stale behavior often hides in derived state, not the primary truth.
- Fix state models, don't mask symptoms: don't paper over problems with bigger limits, buffers, or retries.
- Prefer correcting underlying logic over adding special-case branches. Only special-handle when no cleaner alternative exists.
- If a bug is triggered by entering/leaving/recovering from a state, verify every entry path that reaches the relevant helper, not just the reproduced path.
- After each fix, verify the normal path, failure path, and recovery path.
Exploration
For broad codebase searches (finding all callers of a function, locating cross-module patterns), use Agent(subagent_type: "Explore") instead of chaining Grep/Glob calls.
Deep Reference
This skill's references/ directory contains:
| Reference | Topic | Load When |
|---|---|---|
iterative-debug-case-study.md |
7-round progressive isolation methodology | Stuck after multiple fix attempts; need a structured debugging approach |
Files (embedded-workbench)
-
references
-
iterative-debug-case-study.md 6.8 KB
# Iterative Debugging: A Case Study in Progressive Isolation This document models a real embedded debugging journey — not the specific bug, but the **methodology** that uncovered it across 7 rounds of progressive refinement. Use this as a reference for structuring your own debugging sessions. ## The Pattern: 7 Rounds of Progressive Narrowing ### Round 1 — Symptom: "Loading spinner never resolves" **Initial observation**: After power-on, a sensor display page shows a loading state indefinitely. The sensor appears to be working — data is arriving at the driver level. **First hypothesis**: The warmup timer is too short. **Action**: Double the warmup period. **Result**: No change. The problem is not timing. **Lesson**: Don't tune constants without understanding the state machine. "Not waiting long enough" is the most common wrong first hypothesis. ### Round 2 — Mechanism: Timer vs. Driver ownership **Observation**: The warmup timer lives in the sensor aggregation layer. The sensor driver has its own initialization state flowing independently. **Hypothesis**: The timer and driver initialization are racing — the timer expires before the driver completes init. **Action**: Move the warmup gate from the aggregation layer into the driver, where it can directly observe initialization completion. **Result**: Improves reliability but doesn't fully fix. Some edge cases remain. **Lesson**: Cache freshness ≠ source data readiness. The timer firing means "enough time passed," not "the sensor is ready." Couple the gate to the actual readiness signal. ### Round 3 — Edge case: Black screen after sensor disconnect/reconnect **Symptom**: Unplugging and re-plugging the sensor during operation causes a permanent black screen instead of recovery. **Investigation**: A global animation timer's callback is firing on a freed LVGL object, corrupting the event dispatch chain (HardFault: INVSTATE, LR in event_send_core, PC in SRAM). **Root cause**: The animation timer outlives the page it belongs to. On page exit, the timer is deleted, but a race window allows one final callback to fire on freed memory. **Fix**: Add a generation counter to the animation timer module. Each Create/Delete cycle increments the counter. The callback checks the generation against its captured value and safely returns if mismatched. **Lesson**: Timer lifecycle bugs produce crashes with a distinctive signature: LR in event dispatch, PC in data memory. When you see this, audit all timer Create/Delete pairs before touching any other code. ### Round 4 — Architecture: Gating order matters **Symptom**: The sensor recovers after disconnect, but the display shows "Loading" instead of the measurement. **Investigation**: The recovery path's gating order is wrong. Warmup completion is checked **before** the communication error flag is cleared, so the warmup condition is satisfied (timer expired) while the error condition still blocks display. **Fix**: Reorder the gates: check for errors first, then warmup, then data validity. Each gate must explicitly pass before proceeding to the next. **Lesson**: Multiple independent conditions at a transition gate — verify each one explicitly. Don't assume "timer expired = everything healthy." ### Round 5 — Systemic flaw: One-directional state latch **Symptom**: Once the display reaches "Ready" state, it never returns to "Loading" even when the sensor is disconnected and reconnected. The value briefly shows stale data, then disappears. **Investigation**: The state machine transitions `WarmingUp → Ready` when warmup completes, but has no reverse path. When the sensor later encounters an error, the state stays Ready because no code path resets it. **Root cause**: The state transition model assumes forward-only progress. Real systems need bidirectional transitions. **Fix**: Add a reverse guard at the top of the state update function: if `state == Ready && error_active()`, reset to `WarmingUp`. This runs before any forward transitions. **Lesson**: **One-directional latches are a systemic anti-pattern.** If a target state's preconditions can become false while in that state, you need a reverse transition. Audit every state variable: can it ever need to go backward? ### Round 6 — Ghost data: Derived state not cleared on reset **Symptom**: After the Round 5 fix, a brief flicker shows a stale measurement value before Loading appears. **Investigation**: The state reset (Ready → WarmingUp) clears `display_state` but does NOT clear `cached_sensor_value`. The old valid cached value passes a downstream `if (value > 0) → show Ready` check in the narrow window before the warmup timer starts. **Fix**: In the reset function, clear `cached_sensor_value = NAN` alongside the state reset. Both must be cleared atomically by a single entry point. **Lesson**: **Derived state invariant**: When resetting a primary state, also reset all derived/cached values computed from it. A single stale derived value can bypass every guard in the system. ### Round 7 — Consolidation: Single computation point **Symptom**: Multiple display pages have slightly different loading/error display logic, causing inconsistent behavior across the UI. **Investigation**: The display strategy logic (show value vs. show loading vs. show error) is duplicated across pages, with subtle variations. **Fix**: Extract a single `compute_display_strategy(actual_state, error_info)` function. All pages call it. The function owns the decision; pages only render the result. **Lesson**: When the same decision logic appears in 3+ places, consolidate it. Pages should consume display decisions, not compute them. ## Methodology Summary | Round | What Changed | Method | |-------|-------------|--------| | 1 | Nothing | Tuning constants without understanding — **don't do this** | | 2 | Reliability improved | Moved gate to data source — coupling check to actual signal | | 3 | Crash fixed | Generation counter pattern — timer lifecycle safety | | 4 | Recovery fixed | Gate ordering — explicit precondition verification | | 5 | State fixed | Bidirectional transition — reverse guard pattern | | 6 | Flicker fixed | Derived state invariant — atomic reset | | 7 | Architecture fixed | Single computation point — consolidation | ## Key Takeaways 1. **Start with the state machine, not the symptoms.** Round 1 wasted time on a constant that Round 5 proved irrelevant. 2. **Each fix reveals the next layer.** Don't try to fix everything at once. Round 2 exposed Round 3; Round 5 exposed Round 6. 3. **Derived state is the most common source of subtle bugs.** Cached values, computed flags, display strategies — anything not at the source of truth. 4. **One-directional state latches are always wrong eventually.** If you can't answer "what makes it go back?", you have a bug waiting to happen. 5. **Consolidation happens last, not first.** Fix the individual bugs before extracting common patterns.
-
-
SKILL.md 5.3 KB
--- name: debug-methodology description: "Use when debugging embedded firmware issues — analyzing crash logs, investigating state machine lockups, tracing sensor/signal anomalies, or performing structured root-cause analysis after a crash has been located. For fault-register and stack-frame triage, load hardfault-triage first." --- # Debug Methodology Debug by tracing values, not symptoms. These patterns come from real debugging sessions where surface-level fixes failed and root-cause analysis succeeded. **REQUIRED BACKGROUND:** If the issue involves state machines or protocol timeouts, load `Skill("state-machine-design")` first. If the issue involves FreeRTOS tasks, ISRs, or NVM storage, load `Skill("embedded-firmware-dev")` first. Understand the domain rules before applying debugging methodology. ## Red Flags | You think | Reality | |-----------|---------| | "I know where the bug is, let me fix it" | You know the symptom location. The root cause is often 3 layers away in a different module. | | "One more round of trial fixes and I'll get it" | After 2 failed attempts, you need methodology, not persistence. | | "I'll just add a bounds check and call it done" | You're masking a symptom. The real fix is at the data source, not at every consumer. | | "The logs look normal, it must be a hardware glitch" | If you haven't correlated timestamps to code branches, you haven't actually read the logs. | ## Iron Rules 1. **Log first, not code first**: Correlate serial log timestamps to code branches before touching code. One log line at a known timestamp is worth more than reading five source files. 2. **Call-point census**: When a value isn't updating, find all call sites of the update function. A single call site (e.g., `sensor_update()` only at `main_loop.c:47`) immediately explains why refresh is delayed or gated by irrelevant conditions. ```text grep -rn "sensor_update" --include="*.c" src/main_loop.c:128: sensor_update(&g_pressure); ``` Only two call sites. If `main_loop.c` is gated on `wifi_connected`, the sensor won't refresh until WiFi connects — that's the root cause. 3. **Cache freshness ≠ source data readiness**: In embedded systems, "sensor has data" and "derived cache is refreshed" are concurrent events. Prefer **pull mode** (sync cache on read when source is ready). Avoid pure push mode (periodic background updaters may not fire before the first consumer reads). ```c // BAD: push mode — timer callback pushes data before consumer asks for it static void sensor_timer_cb(TimerHandle_t xTimer) { g_sensor.cache = sensor_read_raw(); // Timer owns refresh timing } float get_temperature(void) { return g_sensor.cache.temperature; // Stale if timer hasn't fired yet } // GOOD: pull mode — consumer triggers refresh when source is ready float get_temperature(void) { if (sensor_is_ready()) { sensor_sync_cache(&g_sensor); // Refresh on read } return g_sensor.cache.temperature; // As fresh as hardware allows } ``` 4. **Multi-path convergence**: When multiple independent code paths show the same error, find the shared state or cache they all read. Fix once at the update point — smaller fix, and future callers can't bypass it. 5. **Ownership boundary mapping**: Before changing behavior, identify which module owns the truth, which derives policy, and which only consumes. Don't let the GUI control backlight policy. Don't let the backlight module control DND state. 6. **Progressive narrowing**: Each investigation round shrinks scope — phenomenon → mechanism → specific state → root cause. Don't try to solve everything at once. 7. **Minimal root-cause fix**: The fix is usually 1-2 lines at the data source. If you're changing 5+ call sites, stop and ask: what single state change would make all of them correct without modification? 8. **Library source is truth**: After 2-3 rounds of custom implementation failure, stop iterating. Read the library source code (e.g., LVGL's `lv_line.c`, `lv_chart.c`) to understand the native mechanism. Adopt and adapt. Verified patterns beat custom math. ## Fix Principles - Draw the failure signature, event timeline, and state transition chain before changing code. - For defects near state transitions: trace the **exact state consumed** by that output. Stale behavior often hides in derived state, not the primary truth. - **Fix state models, don't mask symptoms**: don't paper over problems with bigger limits, buffers, or retries. - Prefer correcting underlying logic over adding special-case branches. Only special-handle when no cleaner alternative exists. - If a bug is triggered by entering/leaving/recovering from a state, **verify every entry path** that reaches the relevant helper, not just the reproduced path. - After each fix, verify the normal path, failure path, and recovery path. ## Exploration For broad codebase searches (finding all callers of a function, locating cross-module patterns), use `Agent(subagent_type: "Explore")` instead of chaining Grep/Glob calls. ## Deep Reference This skill's `references/` directory contains: | Reference | Topic | Load When | |-----------|-------|-----------| | `iterative-debug-case-study.md` | 7-round progressive isolation methodology | Stuck after multiple fix attempts; need a structured debugging approach |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.