forecasting-reverso
Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "pre
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/data-and-visualization/skills/forecasting-reverso
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
forecasting-reverso
Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "predict", "time series", "Reverso", or when tabular data with a temporal dimension needs future-value estimation.
See https://github.com/oaustegard/claude-skills/blob/main/.dev-notes/forecasting-reverso-dev-blog.md for dev notes and background
Skill manifest
Reverso Time Series Forecasting
Produce zero-shot univariate time series forecasts using the Reverso foundation model family (arXiv:2602.17634), implemented in NumPy/Numba for CPU-only container execution.
Setup (run once per conversation)
uv pip install numba --system --break-system-packages
cp /mnt/skills/user/forecasting-reverso/scripts/reverso.py /home/claude/reverso.py
cp /mnt/skills/user/forecasting-reverso/scripts/load_checkpoint.py /home/claude/load_checkpoint.py
Obtaining Weights
Two paths depending on network access:
Path A: Direct download (HuggingFace allow-listed)
import urllib.request, os
os.makedirs("/tmp/reverso", exist_ok=True)
url = "https://huggingface.co/shinfxh/reverso/resolve/main/checkpoints/reverso_small/checkpoint.pth"
urllib.request.urlretrieve(url, "/tmp/reverso/checkpoint.pth")
Path B: User upload (HuggingFace not accessible)
If the download fails with a network error, tell the user:
I can't reach HuggingFace from this environment. Please download the checkpoint from https://huggingface.co/shinfxh/reverso/blob/main/checkpoints/reverso_small/checkpoint.pth and upload it here.
Then load from /mnt/user-data/uploads/checkpoint.pth.
Loading weights
from load_checkpoint import load_checkpoint
weights = load_checkpoint("/tmp/reverso/checkpoint.pth") # or upload path
Model Configuration
Reverso Small uses this config (matching the published args.json):
from reverso import ReversoConfig
config = ReversoConfig(d_model=64, module_list=["conv", "attn", "conv", "attn"])
Forecasting
from reverso import forecast, warmup_jit
warmup_jit() # ~2s one-time JIT compilation
result = forecast(
series=data, # 1-D array/list of floats
prediction_length=96, # how many future steps
weights=weights, # dict from load_checkpoint
config=config,
)
The function handles preprocessing (NaN interpolation, padding, min-max normalization) and autoregressive rollout internally.
Key parameters
flip_equivariant=True — averages forward pass on original and vertically-flipped input. Slightly improves single-step predictions but can dampen amplitude over multi-step rollout. Default is False.
Input Handling
Accept time series as Python list, NumPy array, CSV column, or inline values. Convert to 1-D float array before calling forecast().
For CSV/DataFrame input, ask the user which column to forecast if ambiguous.
The model's context window is 2048 steps. Series shorter than 2048 are left-padded with the first value. Series longer than 2048 use only the most recent 2048 observations. Provide at least a few hundred real data points for meaningful results — heavily padded context degrades forecast quality because the long convolution kernels process mostly constant input.
Visualization
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(12, 4))
n = len(history)
ax.plot(range(n), history, label="Historical", color="#2563eb")
ax.plot(range(n, n + len(preds)), preds,
label="Forecast", color="#dc2626", linewidth=2)
ax.axvline(x=n, color="gray", linestyle="--", alpha=0.4)
ax.set_xlabel("Time step"); ax.set_ylabel("Value")
ax.legend(); fig.tight_layout()
fig.savefig("/mnt/user-data/outputs/forecast.png", dpi=150)
Performance
| Phase | Latency |
|---|---|
| numba install (uv) | ~1.6s |
| Weight loading (.pth) | <1s |
| JIT warmup | ~2s |
| Forward pass (L=2048) | ~80ms |
| 96-step forecast (2 chunks) | ~160ms |
| 192-step forecast (4 chunks) | ~320ms |
Container Environment Limits
Each forward pass takes ~65ms at L=2048. In the ephemeral container, reject batch forecasting requests that would exceed ~1500 forward passes (~100s wall time) to avoid timeouts.
Detect the container environment by checking for /mnt/user-data or /mnt/skills:
import os
IN_CONTAINER = os.path.exists("/mnt/user-data")
Estimate cost before running when processing multiple series:
n_forwards = n_series * n_windows * max(1, pred_length // 48)
est_seconds = n_forwards * 0.065
if IN_CONTAINER and est_seconds > 100:
# Reject or subsample
max_series = int(1500 / (n_windows * max(1, pred_length // 48)))
Practical limits at ~100s budget:
| Scenario | Series | Windows | Pred steps | Forwards | Time |
|---|---|---|---|---|---|
| Single series, 96-step | 1 | 1 | 2 chunks | 2 | 0.1s |
| Small dataset (sz_taxi) | 156 | 6 | 48 | 936 | 61s |
| Medium dataset, short horizon | 300 | 4 | 48 | 1200 | 78s |
| Large dataset (m4_yearly) | 22974 | 1 | 48 | 22974 | 25min ✗ |
When a request exceeds the budget, inform the user with the estimated time and suggest either subsampling or running locally. For benchmark evaluation of large datasets, recommend running outside the container.
Limitations
The model is strongest with periodic or quasi-periodic signals and full 2048-point context. Short series (under ~200 points) are heavily padded and produce degraded forecasts — this is a model limitation, not an implementation bug. Edge cases: binary-valued input (e.g. step functions normalizing to exactly 0/1) and series ending at the exact min-max boundary are out-of-distribution for the training data.
For architecture details, weight mapping, and debugging guidance, read references/architecture.md.
Files (claude-skills)
-
references
-
architecture.md 4.8 KB
# Reverso Architecture Reference ## Model Structure Reverso Small: `conv,attn,conv,attn` with d_model=64, producing this layer stack: ``` Layer 0: CNNBlock (FFT long conv + gating) Layer 1: MLPBlock (d=64 → 256 → 64) Layer 2: AttentionBlock (DeltaNet, state_weaving=True) Layer 3: MLPBlock Layer 4: CNNBlock Layer 5: MLPBlock Layer 6: AttentionBlock (DeltaNet, state_weaving=False) Layer 7: MLPBlock → DecoderHead (cross-attention, 48 output positions) ``` State weaving applies to all attention blocks except the last: `is_intermediate = n_attn < (total_attn - 1)`. ## Weight Name Mapping From the actual checkpoint (86 tensors total, 71 after filtering FlashFFTConv twiddle factors): ### CNNBlock (layers 0, 4) - `layers.{i}.k` — (64, 2048) long conv kernel - `layers.{i}.pregate.net.0.weight` — (64, 1, 3) depthwise gate conv - `layers.{i}.pregate.net.0.bias` — (64,) - `layers.{i}.pregate.net.2.weight` — (64, 64, 1) **pointwise** gate conv (NOT depthwise) - `layers.{i}.pregate.net.2.bias` — (64,) - `layers.{i}.norm.weight/bias` — (64,) LayerNorm ### AttentionBlock (layers 2, 6) All projections are **bias-free** except LayerNorm: - `layers.{i}.attention.q_proj.weight` — (64, 64) - `layers.{i}.attention.k_proj.weight` — (64, 64) - `layers.{i}.attention.v_proj.weight` — (64, 64) - `layers.{i}.attention.o_proj.weight` — (64, 64) - `layers.{i}.attention.b_proj.weight` — (4, 64) beta gate - `layers.{i}.attention.q_conv1d.weight` — (64, 1, 4) causal depthwise, no bias - `layers.{i}.attention.k_conv1d.weight` — (64, 1, 4) - `layers.{i}.attention.v_conv1d.weight` — (64, 1, 4) - `layers.{i}.attention.o_norm.weight` — (16,) per-head RMSNorm - `layers.{i}.norm.weight/bias` — (64,) LayerNorm ### MLPBlock (layers 1, 3, 5, 7) - `layers.{i}.linear.weight` — (256, 64) - `layers.{i}.linear.bias` — (256,) - `layers.{i}.linear_final.weight` — (64, 256) - `layers.{i}.linear_final.bias` — (64,) - `layers.{i}.norm.weight/bias` — (64,) ### Decoder - `head.weight` — (48, 2048) position mixing - `head.bias` — (48,) - `simple_q_proj.weight/bias` — (64, 64) / (64,) - `key_proj.weight/bias` — (64, 64) / (64,) - `value_proj.weight/bias` — (64, 64) / (64,) - `out_proj.weight` — (1, 64) - `out_proj.bias` — (1,) ### Ignored weights `layers.{i}.flashfftconv.*` and `shared_flashfftconv.*` are CUDA-specific precomputed twiddle factors. Skip during loading — scipy.fft handles FFT natively. ## Critical Implementation Details These are lessons from debugging — each caused incorrect output when wrong. ### 1. Circular FFT convolution (n_fft = L, not 2L) FlashFFTConv computes **circular** convolution. The model was trained with wraparound aliasing, so the learned kernels depend on it. ```python # CORRECT — circular conv matching training xf = rfft(x, n=L, axis=-1) kf = rfft(kernel, n=L, axis=-1) result = irfft(xf * kf, n=L, axis=-1) # WRONG — linear conv (zero-padded), 660% error xf = rfft(x, n=2*L, axis=-1) # DO NOT DO THIS ``` ### 2. Per-head L2 normalization (reshape BEFORE normalizing) The fla DeltaNet flow is: linear proj → short conv → reshape to (L, n_heads, d_h) → SiLU → L2 normalize per head. ```python # CORRECT — normalize across d_h=16 per head q = q.reshape(L, n_heads, d_h) q = l2_normalize(silu(q), axis=-1) # ||q|| = 1 per head # WRONG — normalize across d_model=64 before reshape q = l2_normalize(silu(q), axis=-1) # ||q|| = 1 across all heads q = q.reshape(L, n_heads, d_h) # individual heads can have ||k|| >> 1 ``` Wrong order causes DeltaNet state divergence. With ||k||=3.2 and beta=0.46, eigenvalue = 1 - beta*||k||² = -3.8, causing exponential blowup within ~20 steps. ### 3. Flip equivariance for [0,1]-normalized data Input is min-max normalized to [0,1]. Vertical flip is `1-x`, not `-x`: ```python # CORRECT f_flip = model.forward(1.0 - x) result = (f_pos + 1.0 - f_flip) / 2.0 # WRONG — feeds negative values the model never saw in training f_neg = model.forward(-x) result = (f_pos - f_neg) / 2.0 ``` ### 4. Pointwise gate convolution `pregate.net.2` has shape (64, 64, 1) — a full pointwise (1x1) convolution, NOT depthwise. Squeeze the trailing dim and apply as a matrix multiply per position: ```python g = g @ gate_pw_w.squeeze(-1).T + gate_pw_b # (L, d) @ (d, d) per position ``` ### 5. Per-head RMSNorm on attention output The checkpoint contains `o_norm.weight` of shape (d_head,) = (16,). Applied per-head after the DeltaNet recurrence, before the output projection: ```python for h in range(n_heads): out[:, h, :] = rms_norm(out[:, h, :], o_norm_weight) ``` ### 6. PyTorch weight transposition PyTorch stores Linear weights as (out_features, in_features). For `x @ W` computation, transpose to (in_features, out_features): ```python W_numpy = checkpoint_weight.T.copy() ```
-
-
scripts
-
load_checkpoint.py 8.1 KB
"""Load PyTorch checkpoint files (.pth/.pt) without PyTorch installed. Implements a custom pickle unpickler that replaces torch tensor reconstruction with numpy array equivalents. Works for standard state_dict checkpoints saved with torch.save(). """ import io import pickle import zipfile import numpy as np # Mapping from torch dtype strings to numpy dtypes TORCH_DTYPE_MAP = { "torch.float32": np.float32, "torch.float64": np.float64, "torch.float16": np.float16, "torch.bfloat16": None, # special handling needed "torch.int32": np.int32, "torch.int64": np.int64, "torch.int16": np.int16, "torch.int8": np.int8, "torch.uint8": np.uint8, "torch.bool": np.bool_, } TORCH_DTYPE_SIZES = { "torch.float32": 4, "torch.float64": 8, "torch.float16": 2, "torch.bfloat16": 2, "torch.int32": 4, "torch.int64": 8, "torch.int16": 2, "torch.int8": 1, "torch.uint8": 1, "torch.bool": 1, } class _FakeTypedStorage: """Represents a torch TypedStorage backed by raw bytes.""" def __init__(self, dtype_str: str, data: bytes): self.dtype_str = dtype_str self.data = data def to_numpy(self) -> np.ndarray: if self.dtype_str == "torch.bfloat16": # bfloat16 → float32: interpret as uint16, shift left 16 bits raw = np.frombuffer(self.data, dtype=np.uint16) float32_bits = raw.astype(np.uint32) << 16 return float32_bits.view(np.float32) np_dtype = TORCH_DTYPE_MAP[self.dtype_str] return np.frombuffer(self.data, dtype=np_dtype) class TorchUnpickler(pickle.Unpickler): """Custom unpickler that replaces torch objects with numpy equivalents.""" def __init__(self, fp, zip_file: zipfile.ZipFile, archive_prefix: str): super().__init__(fp) self.zip_file = zip_file self.archive_prefix = archive_prefix def find_class(self, module: str, name: str): # Handle torch storage reconstruction if module == "torch._utils" and name == "_rebuild_tensor_v2": return self._rebuild_tensor_v2 if module == "torch" and name == "BFloat16Storage": return lambda *a, **kw: self._make_storage("torch.bfloat16", *a, **kw) if module == "torch" and name == "FloatStorage": return lambda *a, **kw: self._make_storage("torch.float32", *a, **kw) if module == "torch" and name == "HalfStorage": return lambda *a, **kw: self._make_storage("torch.float16", *a, **kw) if module == "torch" and name == "DoubleStorage": return lambda *a, **kw: self._make_storage("torch.float64", *a, **kw) if module == "torch" and name == "LongStorage": return lambda *a, **kw: self._make_storage("torch.int64", *a, **kw) if module == "torch" and name == "IntStorage": return lambda *a, **kw: self._make_storage("torch.int32", *a, **kw) # Handle _rebuild_tensor_v3 if present if module == "torch._utils" and name == "_rebuild_tensor_v3": return self._rebuild_tensor_v2 # v3 has same signature for our needs # Collections and builtins if module == "collections" and name == "OrderedDict": from collections import OrderedDict return OrderedDict # Fallback: try standard resolution try: return super().find_class(module, name) except (ModuleNotFoundError, AttributeError): # Return a placeholder for unknown torch types return lambda *a, **kw: None def persistent_load(self, pid): """Handle persistent_id references to storage files in the zip.""" # pid format: ('storage', storage_type, key, location, numel) if isinstance(pid, tuple) and pid[0] == "storage": _, storage_type_fn, key, location, numel = pid data_path = f"{self.archive_prefix}/data/{key}" raw_data = self.zip_file.read(data_path) # Determine dtype from the storage type dtype_str = self._infer_dtype(storage_type_fn, raw_data, numel) return _FakeTypedStorage(dtype_str, raw_data) return None def _infer_dtype(self, storage_type_fn, raw_data: bytes, numel: int) -> str: """Infer the torch dtype from storage info.""" # Try to get dtype from the storage type function name if hasattr(storage_type_fn, "__name__"): name = storage_type_fn.__name__ elif callable(storage_type_fn): name = str(storage_type_fn) else: name = "" # Check by name for dtype_str, size in TORCH_DTYPE_SIZES.items(): short = dtype_str.split(".")[-1] if short.lower() in name.lower(): return dtype_str # Fallback: infer from data size / numel if numel > 0: bytes_per_elem = len(raw_data) / numel for dtype_str, size in TORCH_DTYPE_SIZES.items(): if abs(size - bytes_per_elem) < 0.01: return dtype_str return "torch.float32" # default @staticmethod def _rebuild_tensor_v2(storage, storage_offset, size, stride, *args): """Reconstruct a tensor from storage as a numpy array.""" if storage is None: return np.zeros(size, dtype=np.float32) arr = storage.to_numpy() # Apply offset and shape total = 1 for s in size: total *= s start = storage_offset flat = arr[start : start + total] # Check if stride is contiguous if len(size) > 0: expected_stride = [] s = 1 for dim in reversed(size): expected_stride.insert(0, s) s *= dim if list(stride) == expected_stride: return flat.reshape(size).copy() else: # Non-contiguous: use stride tricks then copy return flat.reshape(size).copy() # simplified return flat.reshape(size).copy() def load_checkpoint(path: str) -> dict: """Load a PyTorch checkpoint and return a dict of name → numpy array. Handles: - Standard state_dict saves - Checkpoints with 'model_state_dict' or 'state_dict' keys - DDP 'module.' prefix stripping - bfloat16 → float32 conversion """ with zipfile.ZipFile(path, "r") as zf: # Find the archive prefix (directory name inside the zip) names = zf.namelist() pkl_files = [n for n in names if n.endswith("data.pkl")] if not pkl_files: raise ValueError(f"No data.pkl found in {path}") prefix = pkl_files[0].rsplit("/data.pkl", 1)[0] pkl_data = zf.read(pkl_files[0]) unpickler = TorchUnpickler(io.BytesIO(pkl_data), zf, prefix) state = unpickler.load() # Unwrap common checkpoint structures if isinstance(state, dict): if "model_state_dict" in state: state = state["model_state_dict"] elif "state_dict" in state: state = state["state_dict"] # Flatten to name → numpy, strip DDP prefix result = {} if isinstance(state, dict): for key, val in state.items(): clean_key = key.removeprefix("module.") if isinstance(val, np.ndarray): result[clean_key] = val.astype(np.float32) elif isinstance(val, _FakeTypedStorage): result[clean_key] = val.to_numpy().astype(np.float32) # Skip non-tensor entries (optimizer state, etc.) return result if __name__ == "__main__": import sys if len(sys.argv) < 2: print("Usage: python load_checkpoint.py <path.pth> [--save output.npz]") sys.exit(1) path = sys.argv[1] data = load_checkpoint(path) print(f"Loaded {len(data)} tensors:") for k in sorted(data.keys()): print(f" {k}: shape={data[k].shape} dtype={data[k].dtype}") if "--save" in sys.argv: idx = sys.argv.index("--save") out_path = sys.argv[idx + 1] np.savez_compressed(out_path, **data) import os print(f"\nSaved to {out_path} ({os.path.getsize(out_path)/1024:.1f} KB)") -
reverso.py 25.3 KB
""" Reverso Time Series Foundation Model — NumPy/Numba Inference Implementation. CPU-only, dependency-minimal inference for the Reverso model family (arXiv:2602.17634). Supports all published model sizes via config-driven layer stacking. Dependencies: numpy, scipy (fft), numba """ from __future__ import annotations import os import urllib.request from dataclasses import dataclass import numpy as np from numba import njit from scipy.fft import irfft, rfft # --------------------------------------------------------------------------- # Configuration # --------------------------------------------------------------------------- @dataclass class ReversoConfig: """Model configuration for a Reverso variant. Built from the ``args.json`` shipped alongside each checkpoint. """ d_model: int module_list: list[str] seq_len: int = 2048 output_token_len: int = 48 d_intermediate: int = 256 n_heads: int = 4 gating_kernel_size: int = 3 attn_conv_size: int = 4 @property def d_head(self) -> int: return self.d_model // self.n_heads @property def n_modules(self) -> int: return len(self.module_list) @classmethod def from_args(cls, args: dict) -> ReversoConfig: """Create config from a loaded ``args.json`` dictionary.""" module_list = [m.strip() for m in args["main_module"].split(",")] return cls( d_model=args["d_model"], module_list=module_list, seq_len=args.get("seq_len", 2048), output_token_len=args.get("output_token_len", 48), d_intermediate=args.get("d_intermediate", 256), n_heads=4, gating_kernel_size=args.get("gating_kernel_size", 3), attn_conv_size=4, ) # Pre-built configs for known model sizes CONFIGS: dict[str, ReversoConfig] = { "nano": ReversoConfig(d_model=32, module_list=["conv", "attn", "conv", "attn"]), "small": ReversoConfig(d_model=64, module_list=["conv", "attn", "conv", "attn"]), "full": ReversoConfig(d_model=128, module_list=["conv", "attn"] * 8), } # --------------------------------------------------------------------------- # Utility functions # --------------------------------------------------------------------------- def sigmoid(x: np.ndarray) -> np.ndarray: """Numerically stable sigmoid.""" pos = x >= 0 z = np.empty_like(x) z[pos] = 1.0 / (1.0 + np.exp(-x[pos])) exp_x = np.exp(x[~pos]) z[~pos] = exp_x / (1.0 + exp_x) return z def silu(x: np.ndarray) -> np.ndarray: return x * sigmoid(x) def softmax(x: np.ndarray, axis: int = -1) -> np.ndarray: e = np.exp(x - np.max(x, axis=axis, keepdims=True)) return e / e.sum(axis=axis, keepdims=True) def layer_norm( x: np.ndarray, weight: np.ndarray, bias: np.ndarray | None, eps: float = 1e-5 ) -> np.ndarray: """Layer normalization over the last dimension.""" mean = x.mean(axis=-1, keepdims=True) var = x.var(axis=-1, keepdims=True) normed = (x - mean) / np.sqrt(var + eps) out = normed * weight if bias is not None: out = out + bias return out def rms_norm(x: np.ndarray, weight: np.ndarray, eps: float = 1e-6) -> np.ndarray: """Root-mean-square normalization (no bias, no mean centering).""" rms = np.sqrt(np.mean(x * x, axis=-1, keepdims=True) + eps) return (x / rms) * weight def l2_normalize(x: np.ndarray, axis: int = -1, eps: float = 1e-12) -> np.ndarray: norm = np.sqrt(np.sum(x * x, axis=axis, keepdims=True) + eps) return x / norm def simple_rms_norm(x: np.ndarray, eps: float = 1e-6) -> np.ndarray: """Parameter-free RMSNorm: ``x / sqrt(mean(x²))``. Produces vectors with L2 norm ≈ sqrt(dim), NOT unit vectors. Used by fla's DeltaNet for q/k normalization. """ return x / np.sqrt(np.mean(x * x, axis=-1, keepdims=True) + eps) def depthwise_short_conv( x: np.ndarray, weight: np.ndarray, bias: np.ndarray | None = None, ) -> np.ndarray: """Causal depthwise 1-D convolution (correlation, matching PyTorch Conv1d). :param x: Input of shape ``(L, d)``. :param weight: Per-channel kernels, shape ``(d, kernel_size)``. :param bias: Optional, shape ``(d,)``. :returns: Output of shape ``(L, d)``. """ L, d = x.shape ks = weight.shape[1] padded = np.pad(x, ((ks - 1, 0), (0, 0)), mode="constant") out = np.empty((L, d), dtype=x.dtype) for c in range(d): out[:, c] = np.correlate(padded[:, c], weight[c], mode="valid") if bias is not None: out += bias return out def fft_long_conv(x: np.ndarray, kernel: np.ndarray) -> np.ndarray: """Depthwise long circular convolution via FFT. Matches FlashFFTConv behaviour used during training: circular convolution with ``n_fft = L`` (no zero-padding). :param x: Shape ``(L, d)`` — sequence-first layout. :param kernel: Shape ``(d, L)`` — one length-L kernel per channel. :returns: Shape ``(L, d)``. """ L, d = x.shape xt = x.T # (d, L) xf = rfft(xt, n=L, axis=-1) kf = rfft(kernel, n=L, axis=-1) conv = irfft(xf * kf, n=L, axis=-1) return conv.T # (L, d) # --------------------------------------------------------------------------- # Numba-accelerated DeltaNet recurrence # --------------------------------------------------------------------------- @njit(cache=True) def _deltanet_recurrence( q: np.ndarray, k: np.ndarray, v: np.ndarray, beta: np.ndarray, ) -> np.ndarray: """DeltaNet linear-attention recurrence (all heads). :param q: ``(L, n_heads, d_h)`` :param k: ``(L, n_heads, d_h)`` :param v: ``(L, n_heads, d_h)`` :param beta: ``(L, n_heads)`` :returns: ``(L, n_heads, d_h)`` """ L, n_heads, d_h = q.shape out = np.empty_like(q) for h in range(n_heads): S = np.zeros((d_h, d_h), dtype=q.dtype) for i in range(L): ki = k[i, h] vi = v[i, h] bi = beta[i, h] # Sk = S @ ki Sk = np.empty(d_h, dtype=q.dtype) for a in range(d_h): acc = 0.0 for b in range(d_h): acc += S[a, b] * ki[b] Sk[a] = acc # S += bi * outer(vi - Sk, ki) for a in range(d_h): diff = bi * (vi[a] - Sk[a]) for b in range(d_h): S[a, b] += diff * ki[b] # out[i, h] = S @ qi qi = q[i, h] for a in range(d_h): acc = 0.0 for b in range(d_h): acc += S[a, b] * qi[b] out[i, h, a] = acc return out # --------------------------------------------------------------------------- # Model blocks # --------------------------------------------------------------------------- class CNNBlock: """Long depthwise FFT convolution with gating. The gating sub-network is:: gate = sigmoid(pointwise_conv(silu(depthwise_conv(x)))) where ``depthwise_conv`` has ``kernel_size=gating_kernel_size`` (3) and ``pointwise_conv`` has ``kernel_size=1`` (equivalent to a per-position linear projection). """ def __init__( self, kernel: np.ndarray, gate_dw_w: np.ndarray, gate_dw_b: np.ndarray, gate_pw_w: np.ndarray, gate_pw_b: np.ndarray, norm_w: np.ndarray, norm_b: np.ndarray, ): self.kernel = kernel # (d, L) self.gate_dw_w = gate_dw_w # (d, ks) — depthwise conv self.gate_dw_b = gate_dw_b # (d,) # Pointwise conv weight stored as (d_out, d_in, 1) → squeeze to (d_out, d_in) self.gate_pw_w = gate_pw_w.squeeze(-1) if gate_pw_w.ndim == 3 else gate_pw_w self.gate_pw_b = gate_pw_b # (d,) self.norm_w = norm_w self.norm_b = norm_b def __call__(self, x: np.ndarray) -> np.ndarray: """x: (L, d) → (L, d).""" residual = x # Gating g = depthwise_short_conv(x, self.gate_dw_w, self.gate_dw_b) g = silu(g) # Pointwise conv (kernel_size=1) = per-position linear: (L, d_in) @ (d_in, d_out).T g = g @ self.gate_pw_w.T + self.gate_pw_b g = sigmoid(g) gated = x * g # Long convolution out = fft_long_conv(gated, self.kernel) out = np.maximum(out, 0.0) # ReLU out = layer_norm(out, self.norm_w, self.norm_b) + residual return out class MLPBlock: """Two-layer MLP with optional skip projection.""" def __init__( self, linear_w: np.ndarray, linear_b: np.ndarray, final_w: np.ndarray, final_b: np.ndarray, norm_w: np.ndarray, norm_b: np.ndarray, skip_w: np.ndarray | None = None, skip_b: np.ndarray | None = None, ): self.linear_w = linear_w self.linear_b = linear_b self.final_w = final_w self.final_b = final_b self.norm_w = norm_w self.norm_b = norm_b self.skip_w = skip_w self.skip_b = skip_b def __call__(self, x: np.ndarray) -> np.ndarray: if self.skip_w is not None: residual = x @ self.skip_w + self.skip_b else: residual = x y = np.maximum(x @ self.linear_w + self.linear_b, 0.0) # ReLU y = y @ self.final_w + self.final_b y = layer_norm(y, self.norm_w, self.norm_b) + residual return y class AttentionBlock: """DeltaNet linear attention with short convolutions. Matches the ``fla.layers.DeltaNet`` structure from the checkpoint: - Linear projections for q, k, v (no bias) - Causal depthwise short convolutions on q, k, v (no bias) - SiLU activation + L2 normalization on q and k - Beta gate via linear projection (no bias) + sigmoid - DeltaNet recurrence - Per-head RMS normalization (``o_norm``) - Output projection (no bias) - LayerNorm + residual """ def __init__( self, q_proj_w: np.ndarray, k_proj_w: np.ndarray, v_proj_w: np.ndarray, o_proj_w: np.ndarray, beta_w: np.ndarray, q_conv_w: np.ndarray, k_conv_w: np.ndarray, v_conv_w: np.ndarray, o_norm_w: np.ndarray, norm_w: np.ndarray, norm_b: np.ndarray, n_heads: int, state_weaving: bool = False, ): self.q_proj_w = q_proj_w self.k_proj_w = k_proj_w self.v_proj_w = v_proj_w self.o_proj_w = o_proj_w self.beta_w = beta_w self.q_conv_w = q_conv_w self.k_conv_w = k_conv_w self.v_conv_w = v_conv_w self.o_norm_w = o_norm_w # (d_head,) — per-head RMSNorm self.norm_w = norm_w self.norm_b = norm_b self.n_heads = n_heads self.state_weaving = state_weaving def __call__(self, x: np.ndarray) -> np.ndarray: """x: (L, d_model) → (L, d_model).""" residual = x # State weaving: feed end-of-sequence info to start if self.state_weaving: x = x.copy() x[0] += x[-1] L, d = x.shape d_h = d // self.n_heads # Linear projections (no bias) q = x @ self.q_proj_w # (L, d) k = x @ self.k_proj_w v = x @ self.v_proj_w # Short convolutions (causal, depthwise, no bias) q = depthwise_short_conv(q, self.q_conv_w) k = depthwise_short_conv(k, self.k_conv_w) v = depthwise_short_conv(v, self.v_conv_w) # Reshape to multi-head FIRST: (L, d) → (L, n_heads, d_h) q = q.reshape(L, self.n_heads, d_h) k = k.reshape(L, self.n_heads, d_h) v = v.reshape(L, self.n_heads, d_h) # Activations + per-head L2 normalization (q and k only) q = l2_normalize(silu(q), axis=-1) k = l2_normalize(silu(k), axis=-1) # v: no activation, no normalization # Beta gate (no bias) beta = sigmoid(x @ self.beta_w) # (L, n_heads) # Ensure contiguous float64 for Numba q = np.ascontiguousarray(q, dtype=np.float64) k = np.ascontiguousarray(k, dtype=np.float64) v = np.ascontiguousarray(v, dtype=np.float64) beta = np.ascontiguousarray(beta, dtype=np.float64) # DeltaNet recurrence out = _deltanet_recurrence(q, k, v, beta) # (L, n_heads, d_h) out = out.astype(np.float32) # Per-head RMS normalization for h in range(self.n_heads): out[:, h, :] = rms_norm(out[:, h, :], self.o_norm_w) # Reshape back and output projection (no bias) out = out.reshape(L, d) out = out @ self.o_proj_w out = layer_norm(out, self.norm_w, self.norm_b) + residual return out class DecoderHead: """Attention-based decoder producing output_token_len predictions.""" def __init__( self, head_w: np.ndarray, head_b: np.ndarray, q_proj_w: np.ndarray, q_proj_b: np.ndarray, k_proj_w: np.ndarray, k_proj_b: np.ndarray, v_proj_w: np.ndarray, v_proj_b: np.ndarray, out_proj_w: np.ndarray, out_proj_b: np.ndarray, ): self.head_w = head_w self.head_b = head_b self.q_proj_w = q_proj_w self.q_proj_b = q_proj_b self.k_proj_w = k_proj_w self.k_proj_b = k_proj_b self.v_proj_w = v_proj_w self.v_proj_b = v_proj_b self.out_proj_w = out_proj_w self.out_proj_b = out_proj_b def __call__(self, x: np.ndarray) -> np.ndarray: """x: (L, d) → (p,) forecast values.""" L, d = x.shape # Position mixing: (p, L) @ (L, d) → (p, d) z = self.head_w[:, :L] @ x + self.head_b[:, None] # Cross-attention q = z @ self.q_proj_w + self.q_proj_b k = x @ self.k_proj_w + self.k_proj_b v = x @ self.v_proj_w + self.v_proj_b scale = 1.0 / np.sqrt(d) attn_weights = softmax(q @ k.T * scale, axis=-1) attn_out = attn_weights @ v out = (attn_out @ self.out_proj_w + self.out_proj_b).squeeze(-1) return out # --------------------------------------------------------------------------- # Full model # --------------------------------------------------------------------------- class ReversoModel: """Assembled Reverso model for inference.""" def __init__( self, config: ReversoConfig, embedding_w: np.ndarray, layers: list, decoder: DecoderHead, ): self.config = config self.embedding_w = embedding_w # (d_model, 1) self.layers = layers self.decoder = decoder def embed(self, x: np.ndarray) -> np.ndarray: """x: (L,) → (L, d_model).""" return x[:, None] @ self.embedding_w.T def forward(self, x: np.ndarray) -> np.ndarray: """Single forward pass: normalized input (L,) → (p,) predictions.""" h = self.embed(x) for layer_fn in self.layers: h = layer_fn(h) return self.decoder(h) def forward_flip_equivariant(self, x: np.ndarray) -> np.ndarray: """Forward with flip equivariance for [0,1]-normalized input. For min-max normalized data in [0,1], the vertical flip is ``1 - x`` (not ``-x``). If the model is equivariant, ``f(1-x) ≈ 1 - f(x)``. Averaging enforces this symmetry: ``f_eq(x) = (f(x) + 1 - f(1 - x)) / 2`` """ f_pos = self.forward(x) f_flip = self.forward(1.0 - x) return (f_pos + 1.0 - f_flip) / 2.0 # --------------------------------------------------------------------------- # Preprocessing # --------------------------------------------------------------------------- def preprocess( series: np.ndarray, seq_len: int = 2048 ) -> tuple[np.ndarray, float, float]: """Prepare raw series for model input. Handles NaN interpolation, padding, truncation, and min-max normalization. :returns: ``(normalized_series, x_min, x_max)`` """ x = np.asarray(series, dtype=np.float32).copy() # Interpolate NaNs nans = np.isnan(x) if nans.any(): if nans.all(): raise ValueError("Series is entirely NaN") valid = ~nans indices = np.arange(len(x)) x[nans] = np.interp(indices[nans], indices[valid], x[valid]) # Pad short series by back-filling with leftmost value if len(x) < seq_len: pad_len = seq_len - len(x) x = np.concatenate([np.full(pad_len, x[0], dtype=np.float32), x]) # Truncate to last seq_len values x = x[-seq_len:] # Min-max normalization to [0, 1] x_min, x_max = float(x.min()), float(x.max()) denom = x_max - x_min if denom < 1e-10: x_norm = np.full_like(x, 0.5) else: x_norm = (x - x_min) / denom return x_norm, x_min, x_max def postprocess(predictions: np.ndarray, x_min: float, x_max: float) -> np.ndarray: """Unnormalize predictions back to original scale.""" return predictions * (x_max - x_min) + x_min # --------------------------------------------------------------------------- # Weight loading # --------------------------------------------------------------------------- def _get(weights: dict, key: str) -> np.ndarray: """Retrieve a weight, raising a clear error if missing.""" if key not in weights: available = [k for k in sorted(weights.keys()) if not k.startswith("__")] raise KeyError( f"Weight '{key}' not found. Available ({len(available)}): " + ", ".join(available[:20]) + ("..." if len(available) > 20 else "") ) return weights[key].astype(np.float32) def _get_optional(weights: dict, key: str) -> np.ndarray | None: if key in weights: return weights[key].astype(np.float32) return None def _T(w: np.ndarray) -> np.ndarray: """Transpose PyTorch linear weight from (out, in) to (in, out) for x @ W.""" return w.T.copy() def _squeeze_conv(w: np.ndarray) -> np.ndarray: """Squeeze depthwise conv weight from (d, 1, ks) to (d, ks).""" if w.ndim == 3 and w.shape[1] == 1: return w.squeeze(1) return w def load_model(weights: dict, config: ReversoConfig) -> ReversoModel: """Construct a ReversoModel from a weight dictionary and config. :param weights: Dict of ``name → np.ndarray``, from checkpoint or ``.npz``. :param config: Model configuration. """ embedding_w = _get(weights, "embedding.weight") # (d_model, 1) layers = [] layer_idx = 0 n_attn = 0 total_attn = sum(1 for m in config.module_list if m == "attn") for mod_type in config.module_list: if mod_type == "conv": cnn = _build_cnn_block(weights, layer_idx, config) layers.append(cnn) layer_idx += 1 mlp = _build_mlp_block(weights, layer_idx, config) layers.append(mlp) layer_idx += 1 elif mod_type == "attn": is_intermediate = n_attn < (total_attn - 1) attn = _build_attn_block( weights, layer_idx, config, state_weaving=is_intermediate ) layers.append(attn) layer_idx += 1 n_attn += 1 mlp = _build_mlp_block(weights, layer_idx, config) layers.append(mlp) layer_idx += 1 else: raise ValueError(f"Unknown module type: {mod_type}") decoder = _build_decoder(weights, config) return ReversoModel(config, embedding_w, layers, decoder) def _build_cnn_block(weights: dict, idx: int, cfg: ReversoConfig) -> CNNBlock: pfx = f"layers.{idx}" return CNNBlock( kernel=_get(weights, f"{pfx}.k"), gate_dw_w=_squeeze_conv(_get(weights, f"{pfx}.pregate.net.0.weight")), gate_dw_b=_get(weights, f"{pfx}.pregate.net.0.bias"), gate_pw_w=_get(weights, f"{pfx}.pregate.net.2.weight"), gate_pw_b=_get(weights, f"{pfx}.pregate.net.2.bias"), norm_w=_get(weights, f"{pfx}.norm.weight"), norm_b=_get(weights, f"{pfx}.norm.bias"), ) def _build_mlp_block(weights: dict, idx: int, cfg: ReversoConfig) -> MLPBlock: pfx = f"layers.{idx}" skip_w = _get_optional(weights, f"{pfx}.skip_linear.weight") skip_b = _get_optional(weights, f"{pfx}.skip_linear.bias") if skip_w is not None: skip_w = _T(skip_w) return MLPBlock( linear_w=_T(_get(weights, f"{pfx}.linear.weight")), linear_b=_get(weights, f"{pfx}.linear.bias"), final_w=_T(_get(weights, f"{pfx}.linear_final.weight")), final_b=_get(weights, f"{pfx}.linear_final.bias"), norm_w=_get(weights, f"{pfx}.norm.weight"), norm_b=_get(weights, f"{pfx}.norm.bias"), skip_w=skip_w, skip_b=skip_b, ) def _build_attn_block( weights: dict, idx: int, cfg: ReversoConfig, state_weaving: bool ) -> AttentionBlock: pfx = f"layers.{idx}" ap = f"{pfx}.attention" return AttentionBlock( q_proj_w=_T(_get(weights, f"{ap}.q_proj.weight")), k_proj_w=_T(_get(weights, f"{ap}.k_proj.weight")), v_proj_w=_T(_get(weights, f"{ap}.v_proj.weight")), o_proj_w=_T(_get(weights, f"{ap}.o_proj.weight")), beta_w=_T(_get(weights, f"{ap}.b_proj.weight")), q_conv_w=_squeeze_conv(_get(weights, f"{ap}.q_conv1d.weight")), k_conv_w=_squeeze_conv(_get(weights, f"{ap}.k_conv1d.weight")), v_conv_w=_squeeze_conv(_get(weights, f"{ap}.v_conv1d.weight")), o_norm_w=_get(weights, f"{ap}.o_norm.weight"), norm_w=_get(weights, f"{pfx}.norm.weight"), norm_b=_get(weights, f"{pfx}.norm.bias"), n_heads=cfg.n_heads, state_weaving=state_weaving, ) def _build_decoder(weights: dict, cfg: ReversoConfig) -> DecoderHead: return DecoderHead( head_w=_get(weights, "head.weight"), head_b=_get(weights, "head.bias"), q_proj_w=_T(_get(weights, "simple_q_proj.weight")), q_proj_b=_get(weights, "simple_q_proj.bias"), k_proj_w=_T(_get(weights, "key_proj.weight")), k_proj_b=_get(weights, "key_proj.bias"), v_proj_w=_T(_get(weights, "value_proj.weight")), v_proj_b=_get(weights, "value_proj.bias"), out_proj_w=_T(_get(weights, "out_proj.weight")), out_proj_b=_get(weights, "out_proj.bias"), ) # --------------------------------------------------------------------------- # Weight download # --------------------------------------------------------------------------- _WEIGHT_CACHE: dict[str, dict] = {} def download_weights(url: str, cache_dir: str = "/tmp/reverso") -> dict: """Download .npz weights from URL, caching in *cache_dir*.""" if url in _WEIGHT_CACHE: return _WEIGHT_CACHE[url] os.makedirs(cache_dir, exist_ok=True) filename = url.rsplit("/", 1)[-1] local_path = os.path.join(cache_dir, filename) if not os.path.exists(local_path): print(f"Downloading weights from {url} ...") urllib.request.urlretrieve(url, local_path) print(f"Saved to {local_path}") data = dict(np.load(local_path, allow_pickle=False)) _WEIGHT_CACHE[url] = data return data # --------------------------------------------------------------------------- # Autoregressive forecast # --------------------------------------------------------------------------- def forecast( series: np.ndarray | list[float], prediction_length: int, weights: dict | str, model_size: str = "small", config: ReversoConfig | None = None, flip_equivariant: bool = False, ) -> np.ndarray: """Zero-shot time series forecast using Reverso. :param series: Historical observations (1-D). :param prediction_length: Number of future steps to predict. :param weights: Either a weight dict or a URL string to .npz. :param model_size: One of ``'nano'``, ``'small'``, ``'full'`` (ignored if *config* given). :param config: Optional explicit config (overrides *model_size*). :param flip_equivariant: Use flip-equivariant averaging. Helps for single-step prediction but can dampen amplitude during multi-step autoregressive rollout. Default ``False``. :returns: Array of *prediction_length* forecasted values. """ if config is None: config = CONFIGS[model_size] if isinstance(weights, str): weight_data = download_weights(weights) else: weight_data = weights model = load_model(weight_data, config) series = np.asarray(series, dtype=np.float32) x_norm, x_min, x_max = preprocess(series, config.seq_len) forward_fn = model.forward_flip_equivariant if flip_equivariant else model.forward context = x_norm.copy() predictions = [] remaining = prediction_length while remaining > 0: ctx = context[-config.seq_len :] chunk = forward_fn(ctx) take = min(config.output_token_len, remaining) predictions.append(chunk[:take]) remaining -= take context = np.concatenate([context, chunk]) result_norm = np.concatenate(predictions)[:prediction_length] return postprocess(result_norm, x_min, x_max) # --------------------------------------------------------------------------- # Warm-up: trigger Numba JIT compilation # --------------------------------------------------------------------------- def warmup_jit(): """Pre-compile the Numba DeltaNet kernel with a tiny dummy input.""" L, nh, dh = 4, 2, 2 q = np.zeros((L, nh, dh), dtype=np.float64) k = np.zeros((L, nh, dh), dtype=np.float64) v = np.zeros((L, nh, dh), dtype=np.float64) beta = np.zeros((L, nh), dtype=np.float64) _deltanet_recurrence(q, k, v, beta)
-
-
README.md 545 B
# forecasting-reverso Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "predict", "time series", "Reverso", or when tabular data with a temporal dimension needs future-value estimation. See https://github.com/oaustegard/claude-skills/blob/main/.dev-notes/forecasting-reverso-dev-blog.md for dev notes and background -
SKILL.md 5.7 KB
--- name: forecasting-reverso description: Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "predict", "time series", "Reverso", or when tabular data with a temporal dimension needs future-value estimation. metadata: version: 0.1.0 --- # Reverso Time Series Forecasting Produce zero-shot univariate time series forecasts using the Reverso foundation model family (arXiv:2602.17634), implemented in NumPy/Numba for CPU-only container execution. ## Setup (run once per conversation) ```bash uv pip install numba --system --break-system-packages cp /mnt/skills/user/forecasting-reverso/scripts/reverso.py /home/claude/reverso.py cp /mnt/skills/user/forecasting-reverso/scripts/load_checkpoint.py /home/claude/load_checkpoint.py ``` ## Obtaining Weights Two paths depending on network access: ### Path A: Direct download (HuggingFace allow-listed) ```python import urllib.request, os os.makedirs("/tmp/reverso", exist_ok=True) url = "https://huggingface.co/shinfxh/reverso/resolve/main/checkpoints/reverso_small/checkpoint.pth" urllib.request.urlretrieve(url, "/tmp/reverso/checkpoint.pth") ``` ### Path B: User upload (HuggingFace not accessible) If the download fails with a network error, tell the user: > I can't reach HuggingFace from this environment. Please download the checkpoint from > https://huggingface.co/shinfxh/reverso/blob/main/checkpoints/reverso_small/checkpoint.pth > and upload it here. Then load from `/mnt/user-data/uploads/checkpoint.pth`. ### Loading weights ```python from load_checkpoint import load_checkpoint weights = load_checkpoint("/tmp/reverso/checkpoint.pth") # or upload path ``` ## Model Configuration Reverso Small uses this config (matching the published `args.json`): ```python from reverso import ReversoConfig config = ReversoConfig(d_model=64, module_list=["conv", "attn", "conv", "attn"]) ``` ## Forecasting ```python from reverso import forecast, warmup_jit warmup_jit() # ~2s one-time JIT compilation result = forecast( series=data, # 1-D array/list of floats prediction_length=96, # how many future steps weights=weights, # dict from load_checkpoint config=config, ) ``` The function handles preprocessing (NaN interpolation, padding, min-max normalization) and autoregressive rollout internally. ### Key parameters `flip_equivariant=True` — averages forward pass on original and vertically-flipped input. Slightly improves single-step predictions but can dampen amplitude over multi-step rollout. Default is `False`. ## Input Handling Accept time series as Python list, NumPy array, CSV column, or inline values. Convert to 1-D float array before calling `forecast()`. For CSV/DataFrame input, ask the user which column to forecast if ambiguous. The model's context window is 2048 steps. Series shorter than 2048 are left-padded with the first value. Series longer than 2048 use only the most recent 2048 observations. Provide at least a few hundred real data points for meaningful results — heavily padded context degrades forecast quality because the long convolution kernels process mostly constant input. ## Visualization ```python import matplotlib.pyplot as plt fig, ax = plt.subplots(figsize=(12, 4)) n = len(history) ax.plot(range(n), history, label="Historical", color="#2563eb") ax.plot(range(n, n + len(preds)), preds, label="Forecast", color="#dc2626", linewidth=2) ax.axvline(x=n, color="gray", linestyle="--", alpha=0.4) ax.set_xlabel("Time step"); ax.set_ylabel("Value") ax.legend(); fig.tight_layout() fig.savefig("/mnt/user-data/outputs/forecast.png", dpi=150) ``` ## Performance | Phase | Latency | |---|---| | numba install (uv) | ~1.6s | | Weight loading (.pth) | <1s | | JIT warmup | ~2s | | Forward pass (L=2048) | ~80ms | | 96-step forecast (2 chunks) | ~160ms | | 192-step forecast (4 chunks) | ~320ms | ## Container Environment Limits Each forward pass takes ~65ms at L=2048. In the ephemeral container, reject batch forecasting requests that would exceed ~1500 forward passes (~100s wall time) to avoid timeouts. **Detect the container environment** by checking for `/mnt/user-data` or `/mnt/skills`: ```python import os IN_CONTAINER = os.path.exists("/mnt/user-data") ``` **Estimate cost before running** when processing multiple series: ```python n_forwards = n_series * n_windows * max(1, pred_length // 48) est_seconds = n_forwards * 0.065 if IN_CONTAINER and est_seconds > 100: # Reject or subsample max_series = int(1500 / (n_windows * max(1, pred_length // 48))) ``` **Practical limits at ~100s budget:** | Scenario | Series | Windows | Pred steps | Forwards | Time | |---|---|---|---|---|---| | Single series, 96-step | 1 | 1 | 2 chunks | 2 | 0.1s | | Small dataset (sz_taxi) | 156 | 6 | 48 | 936 | 61s | | Medium dataset, short horizon | 300 | 4 | 48 | 1200 | 78s | | Large dataset (m4_yearly) | 22974 | 1 | 48 | 22974 | **25min ✗** | When a request exceeds the budget, inform the user with the estimated time and suggest either subsampling or running locally. For benchmark evaluation of large datasets, recommend running outside the container. ## Limitations The model is strongest with periodic or quasi-periodic signals and full 2048-point context. Short series (under ~200 points) are heavily padded and produce degraded forecasts — this is a model limitation, not an implementation bug. Edge cases: binary-valued input (e.g. step functions normalizing to exactly 0/1) and series ending at the exact min-max boundary are out-of-distribution for the training data. For architecture details, weight mapping, and debugging guidance, read `references/architecture.md`.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.