Claude Skill

forecasting-reverso

Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "pre

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_data-and-visualization_skills_forecasting-reverso-e39c726.zip · 14 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/data-and-visualization/skills/forecasting-reverso
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

forecasting-reverso

Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "predict", "time series", "Reverso", or when tabular data with a temporal dimension needs future-value estimation.

See https://github.com/oaustegard/claude-skills/blob/main/.dev-notes/forecasting-reverso-dev-blog.md for dev notes and background

Skill manifest

Reverso Time Series Forecasting

Produce zero-shot univariate time series forecasts using the Reverso foundation model family (arXiv:2602.17634), implemented in NumPy/Numba for CPU-only container execution.

Setup (run once per conversation)

uv pip install numba --system --break-system-packages
cp /mnt/skills/user/forecasting-reverso/scripts/reverso.py /home/claude/reverso.py
cp /mnt/skills/user/forecasting-reverso/scripts/load_checkpoint.py /home/claude/load_checkpoint.py

Obtaining Weights

Two paths depending on network access:

Path A: Direct download (HuggingFace allow-listed)

import urllib.request, os
os.makedirs("/tmp/reverso", exist_ok=True)
url = "https://huggingface.co/shinfxh/reverso/resolve/main/checkpoints/reverso_small/checkpoint.pth"
urllib.request.urlretrieve(url, "/tmp/reverso/checkpoint.pth")

Path B: User upload (HuggingFace not accessible)

If the download fails with a network error, tell the user:

I can't reach HuggingFace from this environment. Please download the checkpoint from https://huggingface.co/shinfxh/reverso/blob/main/checkpoints/reverso_small/checkpoint.pth and upload it here.

Then load from /mnt/user-data/uploads/checkpoint.pth.

Loading weights

from load_checkpoint import load_checkpoint
weights = load_checkpoint("/tmp/reverso/checkpoint.pth")  # or upload path

Model Configuration

Reverso Small uses this config (matching the published args.json):

from reverso import ReversoConfig
config = ReversoConfig(d_model=64, module_list=["conv", "attn", "conv", "attn"])

Forecasting

from reverso import forecast, warmup_jit
warmup_jit()  # ~2s one-time JIT compilation

result = forecast(
    series=data,               # 1-D array/list of floats
    prediction_length=96,      # how many future steps
    weights=weights,           # dict from load_checkpoint
    config=config,
)

The function handles preprocessing (NaN interpolation, padding, min-max normalization) and autoregressive rollout internally.

Key parameters

flip_equivariant=True — averages forward pass on original and vertically-flipped input. Slightly improves single-step predictions but can dampen amplitude over multi-step rollout. Default is False.

Input Handling

Accept time series as Python list, NumPy array, CSV column, or inline values. Convert to 1-D float array before calling forecast().

For CSV/DataFrame input, ask the user which column to forecast if ambiguous.

The model's context window is 2048 steps. Series shorter than 2048 are left-padded with the first value. Series longer than 2048 use only the most recent 2048 observations. Provide at least a few hundred real data points for meaningful results — heavily padded context degrades forecast quality because the long convolution kernels process mostly constant input.

Visualization

import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(12, 4))
n = len(history)
ax.plot(range(n), history, label="Historical", color="#2563eb")
ax.plot(range(n, n + len(preds)), preds,
        label="Forecast", color="#dc2626", linewidth=2)
ax.axvline(x=n, color="gray", linestyle="--", alpha=0.4)
ax.set_xlabel("Time step"); ax.set_ylabel("Value")
ax.legend(); fig.tight_layout()
fig.savefig("/mnt/user-data/outputs/forecast.png", dpi=150)

Performance

Phase Latency
numba install (uv) ~1.6s
Weight loading (.pth) <1s
JIT warmup ~2s
Forward pass (L=2048) ~80ms
96-step forecast (2 chunks) ~160ms
192-step forecast (4 chunks) ~320ms

Container Environment Limits

Each forward pass takes ~65ms at L=2048. In the ephemeral container, reject batch forecasting requests that would exceed ~1500 forward passes (~100s wall time) to avoid timeouts.

Detect the container environment by checking for /mnt/user-data or /mnt/skills:

import os
IN_CONTAINER = os.path.exists("/mnt/user-data")

Estimate cost before running when processing multiple series:

n_forwards = n_series * n_windows * max(1, pred_length // 48)
est_seconds = n_forwards * 0.065
if IN_CONTAINER and est_seconds > 100:
    # Reject or subsample
    max_series = int(1500 / (n_windows * max(1, pred_length // 48)))

Practical limits at ~100s budget:

Scenario Series Windows Pred steps Forwards Time
Single series, 96-step 1 1 2 chunks 2 0.1s
Small dataset (sz_taxi) 156 6 48 936 61s
Medium dataset, short horizon 300 4 48 1200 78s
Large dataset (m4_yearly) 22974 1 48 22974 25min ✗

When a request exceeds the budget, inform the user with the estimated time and suggest either subsampling or running locally. For benchmark evaluation of large datasets, recommend running outside the container.

Limitations

The model is strongest with periodic or quasi-periodic signals and full 2048-point context. Short series (under ~200 points) are heavily padded and produce degraded forecasts — this is a model limitation, not an implementation bug. Edge cases: binary-valued input (e.g. step functions normalizing to exactly 0/1) and series ending at the exact min-max boundary are out-of-distribution for the training data.

For architecture details, weight mapping, and debugging guidance, read references/architecture.md.

Files (claude-skills)
  • references
    • architecture.md 4.8 KB
      # Reverso Architecture Reference
      
      ## Model Structure
      
      Reverso Small: `conv,attn,conv,attn` with d_model=64, producing this layer stack:
      
      ```
      Layer 0: CNNBlock       (FFT long conv + gating)
      Layer 1: MLPBlock       (d=64 → 256 → 64)
      Layer 2: AttentionBlock (DeltaNet, state_weaving=True)
      Layer 3: MLPBlock
      Layer 4: CNNBlock
      Layer 5: MLPBlock
      Layer 6: AttentionBlock (DeltaNet, state_weaving=False)
      Layer 7: MLPBlock
      → DecoderHead          (cross-attention, 48 output positions)
      ```
      
      State weaving applies to all attention blocks except the last: `is_intermediate = n_attn < (total_attn - 1)`.
      
      ## Weight Name Mapping
      
      From the actual checkpoint (86 tensors total, 71 after filtering FlashFFTConv twiddle factors):
      
      ### CNNBlock (layers 0, 4)
      - `layers.{i}.k` — (64, 2048) long conv kernel
      - `layers.{i}.pregate.net.0.weight` — (64, 1, 3) depthwise gate conv
      - `layers.{i}.pregate.net.0.bias` — (64,)
      - `layers.{i}.pregate.net.2.weight` — (64, 64, 1) **pointwise** gate conv (NOT depthwise)
      - `layers.{i}.pregate.net.2.bias` — (64,)
      - `layers.{i}.norm.weight/bias` — (64,) LayerNorm
      
      ### AttentionBlock (layers 2, 6)
      All projections are **bias-free** except LayerNorm:
      - `layers.{i}.attention.q_proj.weight` — (64, 64)
      - `layers.{i}.attention.k_proj.weight` — (64, 64)
      - `layers.{i}.attention.v_proj.weight` — (64, 64)
      - `layers.{i}.attention.o_proj.weight` — (64, 64)
      - `layers.{i}.attention.b_proj.weight` — (4, 64) beta gate
      - `layers.{i}.attention.q_conv1d.weight` — (64, 1, 4) causal depthwise, no bias
      - `layers.{i}.attention.k_conv1d.weight` — (64, 1, 4)
      - `layers.{i}.attention.v_conv1d.weight` — (64, 1, 4)
      - `layers.{i}.attention.o_norm.weight` — (16,) per-head RMSNorm
      - `layers.{i}.norm.weight/bias` — (64,) LayerNorm
      
      ### MLPBlock (layers 1, 3, 5, 7)
      - `layers.{i}.linear.weight` — (256, 64)
      - `layers.{i}.linear.bias` — (256,)
      - `layers.{i}.linear_final.weight` — (64, 256)
      - `layers.{i}.linear_final.bias` — (64,)
      - `layers.{i}.norm.weight/bias` — (64,)
      
      ### Decoder
      - `head.weight` — (48, 2048) position mixing
      - `head.bias` — (48,)
      - `simple_q_proj.weight/bias` — (64, 64) / (64,)
      - `key_proj.weight/bias` — (64, 64) / (64,)
      - `value_proj.weight/bias` — (64, 64) / (64,)
      - `out_proj.weight` — (1, 64)
      - `out_proj.bias` — (1,)
      
      ### Ignored weights
      `layers.{i}.flashfftconv.*` and `shared_flashfftconv.*` are CUDA-specific precomputed twiddle factors. Skip during loading — scipy.fft handles FFT natively.
      
      ## Critical Implementation Details
      
      These are lessons from debugging — each caused incorrect output when wrong.
      
      ### 1. Circular FFT convolution (n_fft = L, not 2L)
      
      FlashFFTConv computes **circular** convolution. The model was trained with wraparound aliasing, so the learned kernels depend on it.
      
      ```python
      # CORRECT — circular conv matching training
      xf = rfft(x, n=L, axis=-1)
      kf = rfft(kernel, n=L, axis=-1)
      result = irfft(xf * kf, n=L, axis=-1)
      
      # WRONG — linear conv (zero-padded), 660% error
      xf = rfft(x, n=2*L, axis=-1)  # DO NOT DO THIS
      ```
      
      ### 2. Per-head L2 normalization (reshape BEFORE normalizing)
      
      The fla DeltaNet flow is: linear proj → short conv → reshape to (L, n_heads, d_h) → SiLU → L2 normalize per head.
      
      ```python
      # CORRECT — normalize across d_h=16 per head
      q = q.reshape(L, n_heads, d_h)
      q = l2_normalize(silu(q), axis=-1)  # ||q|| = 1 per head
      
      # WRONG — normalize across d_model=64 before reshape
      q = l2_normalize(silu(q), axis=-1)  # ||q|| = 1 across all heads
      q = q.reshape(L, n_heads, d_h)      # individual heads can have ||k|| >> 1
      ```
      
      Wrong order causes DeltaNet state divergence. With ||k||=3.2 and beta=0.46, eigenvalue = 1 - beta*||k||² = -3.8, causing exponential blowup within ~20 steps.
      
      ### 3. Flip equivariance for [0,1]-normalized data
      
      Input is min-max normalized to [0,1]. Vertical flip is `1-x`, not `-x`:
      
      ```python
      # CORRECT
      f_flip = model.forward(1.0 - x)
      result = (f_pos + 1.0 - f_flip) / 2.0
      
      # WRONG — feeds negative values the model never saw in training
      f_neg = model.forward(-x)
      result = (f_pos - f_neg) / 2.0
      ```
      
      ### 4. Pointwise gate convolution
      
      `pregate.net.2` has shape (64, 64, 1) — a full pointwise (1x1) convolution, NOT depthwise. Squeeze the trailing dim and apply as a matrix multiply per position:
      
      ```python
      g = g @ gate_pw_w.squeeze(-1).T + gate_pw_b  # (L, d) @ (d, d) per position
      ```
      
      ### 5. Per-head RMSNorm on attention output
      
      The checkpoint contains `o_norm.weight` of shape (d_head,) = (16,). Applied per-head after the DeltaNet recurrence, before the output projection:
      
      ```python
      for h in range(n_heads):
          out[:, h, :] = rms_norm(out[:, h, :], o_norm_weight)
      ```
      
      ### 6. PyTorch weight transposition
      
      PyTorch stores Linear weights as (out_features, in_features). For `x @ W` computation, transpose to (in_features, out_features):
      
      ```python
      W_numpy = checkpoint_weight.T.copy()
      ```
      
  • scripts
    • load_checkpoint.py 8.1 KB
      """Load PyTorch checkpoint files (.pth/.pt) without PyTorch installed.
      
      Implements a custom pickle unpickler that replaces torch tensor reconstruction
      with numpy array equivalents. Works for standard state_dict checkpoints saved
      with torch.save().
      """
      
      import io
      import pickle
      import zipfile
      
      import numpy as np
      
      # Mapping from torch dtype strings to numpy dtypes
      TORCH_DTYPE_MAP = {
          "torch.float32": np.float32,
          "torch.float64": np.float64,
          "torch.float16": np.float16,
          "torch.bfloat16": None,  # special handling needed
          "torch.int32": np.int32,
          "torch.int64": np.int64,
          "torch.int16": np.int16,
          "torch.int8": np.int8,
          "torch.uint8": np.uint8,
          "torch.bool": np.bool_,
      }
      
      TORCH_DTYPE_SIZES = {
          "torch.float32": 4,
          "torch.float64": 8,
          "torch.float16": 2,
          "torch.bfloat16": 2,
          "torch.int32": 4,
          "torch.int64": 8,
          "torch.int16": 2,
          "torch.int8": 1,
          "torch.uint8": 1,
          "torch.bool": 1,
      }
      
      
      class _FakeTypedStorage:
          """Represents a torch TypedStorage backed by raw bytes."""
      
          def __init__(self, dtype_str: str, data: bytes):
              self.dtype_str = dtype_str
              self.data = data
      
          def to_numpy(self) -> np.ndarray:
              if self.dtype_str == "torch.bfloat16":
                  # bfloat16 → float32: interpret as uint16, shift left 16 bits
                  raw = np.frombuffer(self.data, dtype=np.uint16)
                  float32_bits = raw.astype(np.uint32) << 16
                  return float32_bits.view(np.float32)
              np_dtype = TORCH_DTYPE_MAP[self.dtype_str]
              return np.frombuffer(self.data, dtype=np_dtype)
      
      
      class TorchUnpickler(pickle.Unpickler):
          """Custom unpickler that replaces torch objects with numpy equivalents."""
      
          def __init__(self, fp, zip_file: zipfile.ZipFile, archive_prefix: str):
              super().__init__(fp)
              self.zip_file = zip_file
              self.archive_prefix = archive_prefix
      
          def find_class(self, module: str, name: str):
              # Handle torch storage reconstruction
              if module == "torch._utils" and name == "_rebuild_tensor_v2":
                  return self._rebuild_tensor_v2
      
              if module == "torch" and name == "BFloat16Storage":
                  return lambda *a, **kw: self._make_storage("torch.bfloat16", *a, **kw)
              if module == "torch" and name == "FloatStorage":
                  return lambda *a, **kw: self._make_storage("torch.float32", *a, **kw)
              if module == "torch" and name == "HalfStorage":
                  return lambda *a, **kw: self._make_storage("torch.float16", *a, **kw)
              if module == "torch" and name == "DoubleStorage":
                  return lambda *a, **kw: self._make_storage("torch.float64", *a, **kw)
              if module == "torch" and name == "LongStorage":
                  return lambda *a, **kw: self._make_storage("torch.int64", *a, **kw)
              if module == "torch" and name == "IntStorage":
                  return lambda *a, **kw: self._make_storage("torch.int32", *a, **kw)
      
              # Handle _rebuild_tensor_v3 if present
              if module == "torch._utils" and name == "_rebuild_tensor_v3":
                  return self._rebuild_tensor_v2  # v3 has same signature for our needs
      
              # Collections and builtins
              if module == "collections" and name == "OrderedDict":
                  from collections import OrderedDict
                  return OrderedDict
      
              # Fallback: try standard resolution
              try:
                  return super().find_class(module, name)
              except (ModuleNotFoundError, AttributeError):
                  # Return a placeholder for unknown torch types
                  return lambda *a, **kw: None
      
          def persistent_load(self, pid):
              """Handle persistent_id references to storage files in the zip."""
              # pid format: ('storage', storage_type, key, location, numel)
              if isinstance(pid, tuple) and pid[0] == "storage":
                  _, storage_type_fn, key, location, numel = pid
                  data_path = f"{self.archive_prefix}/data/{key}"
                  raw_data = self.zip_file.read(data_path)
                  # Determine dtype from the storage type
                  dtype_str = self._infer_dtype(storage_type_fn, raw_data, numel)
                  return _FakeTypedStorage(dtype_str, raw_data)
              return None
      
          def _infer_dtype(self, storage_type_fn, raw_data: bytes, numel: int) -> str:
              """Infer the torch dtype from storage info."""
              # Try to get dtype from the storage type function name
              if hasattr(storage_type_fn, "__name__"):
                  name = storage_type_fn.__name__
              elif callable(storage_type_fn):
                  name = str(storage_type_fn)
              else:
                  name = ""
      
              # Check by name
              for dtype_str, size in TORCH_DTYPE_SIZES.items():
                  short = dtype_str.split(".")[-1]
                  if short.lower() in name.lower():
                      return dtype_str
      
              # Fallback: infer from data size / numel
              if numel > 0:
                  bytes_per_elem = len(raw_data) / numel
                  for dtype_str, size in TORCH_DTYPE_SIZES.items():
                      if abs(size - bytes_per_elem) < 0.01:
                          return dtype_str
      
              return "torch.float32"  # default
      
          @staticmethod
          def _rebuild_tensor_v2(storage, storage_offset, size, stride, *args):
              """Reconstruct a tensor from storage as a numpy array."""
              if storage is None:
                  return np.zeros(size, dtype=np.float32)
      
              arr = storage.to_numpy()
              # Apply offset and shape
              total = 1
              for s in size:
                  total *= s
              start = storage_offset
              flat = arr[start : start + total]
      
              # Check if stride is contiguous
              if len(size) > 0:
                  expected_stride = []
                  s = 1
                  for dim in reversed(size):
                      expected_stride.insert(0, s)
                      s *= dim
                  if list(stride) == expected_stride:
                      return flat.reshape(size).copy()
                  else:
                      # Non-contiguous: use stride tricks then copy
                      return flat.reshape(size).copy()  # simplified
      
              return flat.reshape(size).copy()
      
      
      def load_checkpoint(path: str) -> dict:
          """Load a PyTorch checkpoint and return a dict of name → numpy array.
      
          Handles:
          - Standard state_dict saves
          - Checkpoints with 'model_state_dict' or 'state_dict' keys
          - DDP 'module.' prefix stripping
          - bfloat16 → float32 conversion
          """
          with zipfile.ZipFile(path, "r") as zf:
              # Find the archive prefix (directory name inside the zip)
              names = zf.namelist()
              pkl_files = [n for n in names if n.endswith("data.pkl")]
              if not pkl_files:
                  raise ValueError(f"No data.pkl found in {path}")
      
              prefix = pkl_files[0].rsplit("/data.pkl", 1)[0]
      
              pkl_data = zf.read(pkl_files[0])
              unpickler = TorchUnpickler(io.BytesIO(pkl_data), zf, prefix)
              state = unpickler.load()
      
          # Unwrap common checkpoint structures
          if isinstance(state, dict):
              if "model_state_dict" in state:
                  state = state["model_state_dict"]
              elif "state_dict" in state:
                  state = state["state_dict"]
      
          # Flatten to name → numpy, strip DDP prefix
          result = {}
          if isinstance(state, dict):
              for key, val in state.items():
                  clean_key = key.removeprefix("module.")
                  if isinstance(val, np.ndarray):
                      result[clean_key] = val.astype(np.float32)
                  elif isinstance(val, _FakeTypedStorage):
                      result[clean_key] = val.to_numpy().astype(np.float32)
                  # Skip non-tensor entries (optimizer state, etc.)
      
          return result
      
      
      if __name__ == "__main__":
          import sys
          if len(sys.argv) < 2:
              print("Usage: python load_checkpoint.py <path.pth> [--save output.npz]")
              sys.exit(1)
      
          path = sys.argv[1]
          data = load_checkpoint(path)
          print(f"Loaded {len(data)} tensors:")
          for k in sorted(data.keys()):
              print(f"  {k}: shape={data[k].shape} dtype={data[k].dtype}")
      
          if "--save" in sys.argv:
              idx = sys.argv.index("--save")
              out_path = sys.argv[idx + 1]
              np.savez_compressed(out_path, **data)
              import os
              print(f"\nSaved to {out_path} ({os.path.getsize(out_path)/1024:.1f} KB)")
      
    • reverso.py 25.3 KB
      """
      Reverso Time Series Foundation Model — NumPy/Numba Inference Implementation.
      
      CPU-only, dependency-minimal inference for the Reverso model family
      (arXiv:2602.17634). Supports all published model sizes via config-driven
      layer stacking.
      
      Dependencies: numpy, scipy (fft), numba
      """
      
      from __future__ import annotations
      
      import os
      import urllib.request
      from dataclasses import dataclass
      
      import numpy as np
      from numba import njit
      from scipy.fft import irfft, rfft
      
      # ---------------------------------------------------------------------------
      # Configuration
      # ---------------------------------------------------------------------------
      
      
      @dataclass
      class ReversoConfig:
          """Model configuration for a Reverso variant.
      
          Built from the ``args.json`` shipped alongside each checkpoint.
          """
      
          d_model: int
          module_list: list[str]
          seq_len: int = 2048
          output_token_len: int = 48
          d_intermediate: int = 256
          n_heads: int = 4
          gating_kernel_size: int = 3
          attn_conv_size: int = 4
      
          @property
          def d_head(self) -> int:
              return self.d_model // self.n_heads
      
          @property
          def n_modules(self) -> int:
              return len(self.module_list)
      
          @classmethod
          def from_args(cls, args: dict) -> ReversoConfig:
              """Create config from a loaded ``args.json`` dictionary."""
              module_list = [m.strip() for m in args["main_module"].split(",")]
              return cls(
                  d_model=args["d_model"],
                  module_list=module_list,
                  seq_len=args.get("seq_len", 2048),
                  output_token_len=args.get("output_token_len", 48),
                  d_intermediate=args.get("d_intermediate", 256),
                  n_heads=4,
                  gating_kernel_size=args.get("gating_kernel_size", 3),
                  attn_conv_size=4,
              )
      
      
      # Pre-built configs for known model sizes
      CONFIGS: dict[str, ReversoConfig] = {
          "nano": ReversoConfig(d_model=32, module_list=["conv", "attn", "conv", "attn"]),
          "small": ReversoConfig(d_model=64, module_list=["conv", "attn", "conv", "attn"]),
          "full": ReversoConfig(d_model=128, module_list=["conv", "attn"] * 8),
      }
      
      
      # ---------------------------------------------------------------------------
      # Utility functions
      # ---------------------------------------------------------------------------
      
      
      def sigmoid(x: np.ndarray) -> np.ndarray:
          """Numerically stable sigmoid."""
          pos = x >= 0
          z = np.empty_like(x)
          z[pos] = 1.0 / (1.0 + np.exp(-x[pos]))
          exp_x = np.exp(x[~pos])
          z[~pos] = exp_x / (1.0 + exp_x)
          return z
      
      
      def silu(x: np.ndarray) -> np.ndarray:
          return x * sigmoid(x)
      
      
      def softmax(x: np.ndarray, axis: int = -1) -> np.ndarray:
          e = np.exp(x - np.max(x, axis=axis, keepdims=True))
          return e / e.sum(axis=axis, keepdims=True)
      
      
      def layer_norm(
          x: np.ndarray, weight: np.ndarray, bias: np.ndarray | None, eps: float = 1e-5
      ) -> np.ndarray:
          """Layer normalization over the last dimension."""
          mean = x.mean(axis=-1, keepdims=True)
          var = x.var(axis=-1, keepdims=True)
          normed = (x - mean) / np.sqrt(var + eps)
          out = normed * weight
          if bias is not None:
              out = out + bias
          return out
      
      
      def rms_norm(x: np.ndarray, weight: np.ndarray, eps: float = 1e-6) -> np.ndarray:
          """Root-mean-square normalization (no bias, no mean centering)."""
          rms = np.sqrt(np.mean(x * x, axis=-1, keepdims=True) + eps)
          return (x / rms) * weight
      
      
      def l2_normalize(x: np.ndarray, axis: int = -1, eps: float = 1e-12) -> np.ndarray:
          norm = np.sqrt(np.sum(x * x, axis=axis, keepdims=True) + eps)
          return x / norm
      
      
      def simple_rms_norm(x: np.ndarray, eps: float = 1e-6) -> np.ndarray:
          """Parameter-free RMSNorm: ``x / sqrt(mean(x²))``.
      
          Produces vectors with L2 norm ≈ sqrt(dim), NOT unit vectors.
          Used by fla's DeltaNet for q/k normalization.
          """
          return x / np.sqrt(np.mean(x * x, axis=-1, keepdims=True) + eps)
      
      
      def depthwise_short_conv(
          x: np.ndarray,
          weight: np.ndarray,
          bias: np.ndarray | None = None,
      ) -> np.ndarray:
          """Causal depthwise 1-D convolution (correlation, matching PyTorch Conv1d).
      
          :param x: Input of shape ``(L, d)``.
          :param weight: Per-channel kernels, shape ``(d, kernel_size)``.
          :param bias: Optional, shape ``(d,)``.
          :returns: Output of shape ``(L, d)``.
          """
          L, d = x.shape
          ks = weight.shape[1]
          padded = np.pad(x, ((ks - 1, 0), (0, 0)), mode="constant")
          out = np.empty((L, d), dtype=x.dtype)
          for c in range(d):
              out[:, c] = np.correlate(padded[:, c], weight[c], mode="valid")
          if bias is not None:
              out += bias
          return out
      
      
      def fft_long_conv(x: np.ndarray, kernel: np.ndarray) -> np.ndarray:
          """Depthwise long circular convolution via FFT.
      
          Matches FlashFFTConv behaviour used during training: circular convolution
          with ``n_fft = L`` (no zero-padding).
      
          :param x: Shape ``(L, d)`` — sequence-first layout.
          :param kernel: Shape ``(d, L)`` — one length-L kernel per channel.
          :returns: Shape ``(L, d)``.
          """
          L, d = x.shape
          xt = x.T  # (d, L)
          xf = rfft(xt, n=L, axis=-1)
          kf = rfft(kernel, n=L, axis=-1)
          conv = irfft(xf * kf, n=L, axis=-1)
          return conv.T  # (L, d)
      
      
      # ---------------------------------------------------------------------------
      # Numba-accelerated DeltaNet recurrence
      # ---------------------------------------------------------------------------
      
      
      @njit(cache=True)
      def _deltanet_recurrence(
          q: np.ndarray,
          k: np.ndarray,
          v: np.ndarray,
          beta: np.ndarray,
      ) -> np.ndarray:
          """DeltaNet linear-attention recurrence (all heads).
      
          :param q: ``(L, n_heads, d_h)``
          :param k: ``(L, n_heads, d_h)``
          :param v: ``(L, n_heads, d_h)``
          :param beta: ``(L, n_heads)``
          :returns: ``(L, n_heads, d_h)``
          """
          L, n_heads, d_h = q.shape
          out = np.empty_like(q)
      
          for h in range(n_heads):
              S = np.zeros((d_h, d_h), dtype=q.dtype)
              for i in range(L):
                  ki = k[i, h]
                  vi = v[i, h]
                  bi = beta[i, h]
      
                  # Sk = S @ ki
                  Sk = np.empty(d_h, dtype=q.dtype)
                  for a in range(d_h):
                      acc = 0.0
                      for b in range(d_h):
                          acc += S[a, b] * ki[b]
                      Sk[a] = acc
      
                  # S += bi * outer(vi - Sk, ki)
                  for a in range(d_h):
                      diff = bi * (vi[a] - Sk[a])
                      for b in range(d_h):
                          S[a, b] += diff * ki[b]
      
                  # out[i, h] = S @ qi
                  qi = q[i, h]
                  for a in range(d_h):
                      acc = 0.0
                      for b in range(d_h):
                          acc += S[a, b] * qi[b]
                      out[i, h, a] = acc
      
          return out
      
      
      # ---------------------------------------------------------------------------
      # Model blocks
      # ---------------------------------------------------------------------------
      
      
      class CNNBlock:
          """Long depthwise FFT convolution with gating.
      
          The gating sub-network is::
      
              gate = sigmoid(pointwise_conv(silu(depthwise_conv(x))))
      
          where ``depthwise_conv`` has ``kernel_size=gating_kernel_size`` (3) and
          ``pointwise_conv`` has ``kernel_size=1`` (equivalent to a per-position
          linear projection).
          """
      
          def __init__(
              self,
              kernel: np.ndarray,
              gate_dw_w: np.ndarray,
              gate_dw_b: np.ndarray,
              gate_pw_w: np.ndarray,
              gate_pw_b: np.ndarray,
              norm_w: np.ndarray,
              norm_b: np.ndarray,
          ):
              self.kernel = kernel            # (d, L)
              self.gate_dw_w = gate_dw_w      # (d, ks) — depthwise conv
              self.gate_dw_b = gate_dw_b      # (d,)
              # Pointwise conv weight stored as (d_out, d_in, 1) → squeeze to (d_out, d_in)
              self.gate_pw_w = gate_pw_w.squeeze(-1) if gate_pw_w.ndim == 3 else gate_pw_w
              self.gate_pw_b = gate_pw_b      # (d,)
              self.norm_w = norm_w
              self.norm_b = norm_b
      
          def __call__(self, x: np.ndarray) -> np.ndarray:
              """x: (L, d) → (L, d)."""
              residual = x
              # Gating
              g = depthwise_short_conv(x, self.gate_dw_w, self.gate_dw_b)
              g = silu(g)
              # Pointwise conv (kernel_size=1) = per-position linear: (L, d_in) @ (d_in, d_out).T
              g = g @ self.gate_pw_w.T + self.gate_pw_b
              g = sigmoid(g)
              gated = x * g
              # Long convolution
              out = fft_long_conv(gated, self.kernel)
              out = np.maximum(out, 0.0)  # ReLU
              out = layer_norm(out, self.norm_w, self.norm_b) + residual
              return out
      
      
      class MLPBlock:
          """Two-layer MLP with optional skip projection."""
      
          def __init__(
              self,
              linear_w: np.ndarray,
              linear_b: np.ndarray,
              final_w: np.ndarray,
              final_b: np.ndarray,
              norm_w: np.ndarray,
              norm_b: np.ndarray,
              skip_w: np.ndarray | None = None,
              skip_b: np.ndarray | None = None,
          ):
              self.linear_w = linear_w
              self.linear_b = linear_b
              self.final_w = final_w
              self.final_b = final_b
              self.norm_w = norm_w
              self.norm_b = norm_b
              self.skip_w = skip_w
              self.skip_b = skip_b
      
          def __call__(self, x: np.ndarray) -> np.ndarray:
              if self.skip_w is not None:
                  residual = x @ self.skip_w + self.skip_b
              else:
                  residual = x
              y = np.maximum(x @ self.linear_w + self.linear_b, 0.0)  # ReLU
              y = y @ self.final_w + self.final_b
              y = layer_norm(y, self.norm_w, self.norm_b) + residual
              return y
      
      
      class AttentionBlock:
          """DeltaNet linear attention with short convolutions.
      
          Matches the ``fla.layers.DeltaNet`` structure from the checkpoint:
      
          - Linear projections for q, k, v (no bias)
          - Causal depthwise short convolutions on q, k, v (no bias)
          - SiLU activation + L2 normalization on q and k
          - Beta gate via linear projection (no bias) + sigmoid
          - DeltaNet recurrence
          - Per-head RMS normalization (``o_norm``)
          - Output projection (no bias)
          - LayerNorm + residual
          """
      
          def __init__(
              self,
              q_proj_w: np.ndarray,
              k_proj_w: np.ndarray,
              v_proj_w: np.ndarray,
              o_proj_w: np.ndarray,
              beta_w: np.ndarray,
              q_conv_w: np.ndarray,
              k_conv_w: np.ndarray,
              v_conv_w: np.ndarray,
              o_norm_w: np.ndarray,
              norm_w: np.ndarray,
              norm_b: np.ndarray,
              n_heads: int,
              state_weaving: bool = False,
          ):
              self.q_proj_w = q_proj_w
              self.k_proj_w = k_proj_w
              self.v_proj_w = v_proj_w
              self.o_proj_w = o_proj_w
              self.beta_w = beta_w
              self.q_conv_w = q_conv_w
              self.k_conv_w = k_conv_w
              self.v_conv_w = v_conv_w
              self.o_norm_w = o_norm_w    # (d_head,) — per-head RMSNorm
              self.norm_w = norm_w
              self.norm_b = norm_b
              self.n_heads = n_heads
              self.state_weaving = state_weaving
      
          def __call__(self, x: np.ndarray) -> np.ndarray:
              """x: (L, d_model) → (L, d_model)."""
              residual = x
      
              # State weaving: feed end-of-sequence info to start
              if self.state_weaving:
                  x = x.copy()
                  x[0] += x[-1]
      
              L, d = x.shape
              d_h = d // self.n_heads
      
              # Linear projections (no bias)
              q = x @ self.q_proj_w      # (L, d)
              k = x @ self.k_proj_w
              v = x @ self.v_proj_w
      
              # Short convolutions (causal, depthwise, no bias)
              q = depthwise_short_conv(q, self.q_conv_w)
              k = depthwise_short_conv(k, self.k_conv_w)
              v = depthwise_short_conv(v, self.v_conv_w)
      
              # Reshape to multi-head FIRST: (L, d) → (L, n_heads, d_h)
              q = q.reshape(L, self.n_heads, d_h)
              k = k.reshape(L, self.n_heads, d_h)
              v = v.reshape(L, self.n_heads, d_h)
      
              # Activations + per-head L2 normalization (q and k only)
              q = l2_normalize(silu(q), axis=-1)
              k = l2_normalize(silu(k), axis=-1)
              # v: no activation, no normalization
      
              # Beta gate (no bias)
              beta = sigmoid(x @ self.beta_w)  # (L, n_heads)
      
              # Ensure contiguous float64 for Numba
              q = np.ascontiguousarray(q, dtype=np.float64)
              k = np.ascontiguousarray(k, dtype=np.float64)
              v = np.ascontiguousarray(v, dtype=np.float64)
              beta = np.ascontiguousarray(beta, dtype=np.float64)
      
              # DeltaNet recurrence
              out = _deltanet_recurrence(q, k, v, beta)  # (L, n_heads, d_h)
              out = out.astype(np.float32)
      
              # Per-head RMS normalization
              for h in range(self.n_heads):
                  out[:, h, :] = rms_norm(out[:, h, :], self.o_norm_w)
      
              # Reshape back and output projection (no bias)
              out = out.reshape(L, d)
              out = out @ self.o_proj_w
      
              out = layer_norm(out, self.norm_w, self.norm_b) + residual
              return out
      
      
      class DecoderHead:
          """Attention-based decoder producing output_token_len predictions."""
      
          def __init__(
              self,
              head_w: np.ndarray,
              head_b: np.ndarray,
              q_proj_w: np.ndarray,
              q_proj_b: np.ndarray,
              k_proj_w: np.ndarray,
              k_proj_b: np.ndarray,
              v_proj_w: np.ndarray,
              v_proj_b: np.ndarray,
              out_proj_w: np.ndarray,
              out_proj_b: np.ndarray,
          ):
              self.head_w = head_w
              self.head_b = head_b
              self.q_proj_w = q_proj_w
              self.q_proj_b = q_proj_b
              self.k_proj_w = k_proj_w
              self.k_proj_b = k_proj_b
              self.v_proj_w = v_proj_w
              self.v_proj_b = v_proj_b
              self.out_proj_w = out_proj_w
              self.out_proj_b = out_proj_b
      
          def __call__(self, x: np.ndarray) -> np.ndarray:
              """x: (L, d) → (p,) forecast values."""
              L, d = x.shape
              # Position mixing: (p, L) @ (L, d) → (p, d)
              z = self.head_w[:, :L] @ x + self.head_b[:, None]
      
              # Cross-attention
              q = z @ self.q_proj_w + self.q_proj_b
              k = x @ self.k_proj_w + self.k_proj_b
              v = x @ self.v_proj_w + self.v_proj_b
      
              scale = 1.0 / np.sqrt(d)
              attn_weights = softmax(q @ k.T * scale, axis=-1)
              attn_out = attn_weights @ v
      
              out = (attn_out @ self.out_proj_w + self.out_proj_b).squeeze(-1)
              return out
      
      
      # ---------------------------------------------------------------------------
      # Full model
      # ---------------------------------------------------------------------------
      
      
      class ReversoModel:
          """Assembled Reverso model for inference."""
      
          def __init__(
              self,
              config: ReversoConfig,
              embedding_w: np.ndarray,
              layers: list,
              decoder: DecoderHead,
          ):
              self.config = config
              self.embedding_w = embedding_w  # (d_model, 1)
              self.layers = layers
              self.decoder = decoder
      
          def embed(self, x: np.ndarray) -> np.ndarray:
              """x: (L,) → (L, d_model)."""
              return x[:, None] @ self.embedding_w.T
      
          def forward(self, x: np.ndarray) -> np.ndarray:
              """Single forward pass: normalized input (L,) → (p,) predictions."""
              h = self.embed(x)
              for layer_fn in self.layers:
                  h = layer_fn(h)
              return self.decoder(h)
      
          def forward_flip_equivariant(self, x: np.ndarray) -> np.ndarray:
              """Forward with flip equivariance for [0,1]-normalized input.
      
              For min-max normalized data in [0,1], the vertical flip is ``1 - x``
              (not ``-x``).  If the model is equivariant, ``f(1-x) ≈ 1 - f(x)``.
              Averaging enforces this symmetry:
      
                  ``f_eq(x) = (f(x) + 1 - f(1 - x)) / 2``
              """
              f_pos = self.forward(x)
              f_flip = self.forward(1.0 - x)
              return (f_pos + 1.0 - f_flip) / 2.0
      
      
      # ---------------------------------------------------------------------------
      # Preprocessing
      # ---------------------------------------------------------------------------
      
      
      def preprocess(
          series: np.ndarray, seq_len: int = 2048
      ) -> tuple[np.ndarray, float, float]:
          """Prepare raw series for model input.
      
          Handles NaN interpolation, padding, truncation, and min-max normalization.
      
          :returns: ``(normalized_series, x_min, x_max)``
          """
          x = np.asarray(series, dtype=np.float32).copy()
      
          # Interpolate NaNs
          nans = np.isnan(x)
          if nans.any():
              if nans.all():
                  raise ValueError("Series is entirely NaN")
              valid = ~nans
              indices = np.arange(len(x))
              x[nans] = np.interp(indices[nans], indices[valid], x[valid])
      
          # Pad short series by back-filling with leftmost value
          if len(x) < seq_len:
              pad_len = seq_len - len(x)
              x = np.concatenate([np.full(pad_len, x[0], dtype=np.float32), x])
      
          # Truncate to last seq_len values
          x = x[-seq_len:]
      
          # Min-max normalization to [0, 1]
          x_min, x_max = float(x.min()), float(x.max())
          denom = x_max - x_min
          if denom < 1e-10:
              x_norm = np.full_like(x, 0.5)
          else:
              x_norm = (x - x_min) / denom
      
          return x_norm, x_min, x_max
      
      
      def postprocess(predictions: np.ndarray, x_min: float, x_max: float) -> np.ndarray:
          """Unnormalize predictions back to original scale."""
          return predictions * (x_max - x_min) + x_min
      
      
      # ---------------------------------------------------------------------------
      # Weight loading
      # ---------------------------------------------------------------------------
      
      
      def _get(weights: dict, key: str) -> np.ndarray:
          """Retrieve a weight, raising a clear error if missing."""
          if key not in weights:
              available = [k for k in sorted(weights.keys()) if not k.startswith("__")]
              raise KeyError(
                  f"Weight '{key}' not found. Available ({len(available)}): "
                  + ", ".join(available[:20])
                  + ("..." if len(available) > 20 else "")
              )
          return weights[key].astype(np.float32)
      
      
      def _get_optional(weights: dict, key: str) -> np.ndarray | None:
          if key in weights:
              return weights[key].astype(np.float32)
          return None
      
      
      def _T(w: np.ndarray) -> np.ndarray:
          """Transpose PyTorch linear weight from (out, in) to (in, out) for x @ W."""
          return w.T.copy()
      
      
      def _squeeze_conv(w: np.ndarray) -> np.ndarray:
          """Squeeze depthwise conv weight from (d, 1, ks) to (d, ks)."""
          if w.ndim == 3 and w.shape[1] == 1:
              return w.squeeze(1)
          return w
      
      
      def load_model(weights: dict, config: ReversoConfig) -> ReversoModel:
          """Construct a ReversoModel from a weight dictionary and config.
      
          :param weights: Dict of ``name → np.ndarray``, from checkpoint or ``.npz``.
          :param config: Model configuration.
          """
          embedding_w = _get(weights, "embedding.weight")  # (d_model, 1)
      
          layers = []
          layer_idx = 0
      
          n_attn = 0
          total_attn = sum(1 for m in config.module_list if m == "attn")
      
          for mod_type in config.module_list:
              if mod_type == "conv":
                  cnn = _build_cnn_block(weights, layer_idx, config)
                  layers.append(cnn)
                  layer_idx += 1
      
                  mlp = _build_mlp_block(weights, layer_idx, config)
                  layers.append(mlp)
                  layer_idx += 1
      
              elif mod_type == "attn":
                  is_intermediate = n_attn < (total_attn - 1)
                  attn = _build_attn_block(
                      weights, layer_idx, config, state_weaving=is_intermediate
                  )
                  layers.append(attn)
                  layer_idx += 1
                  n_attn += 1
      
                  mlp = _build_mlp_block(weights, layer_idx, config)
                  layers.append(mlp)
                  layer_idx += 1
              else:
                  raise ValueError(f"Unknown module type: {mod_type}")
      
          decoder = _build_decoder(weights, config)
          return ReversoModel(config, embedding_w, layers, decoder)
      
      
      def _build_cnn_block(weights: dict, idx: int, cfg: ReversoConfig) -> CNNBlock:
          pfx = f"layers.{idx}"
          return CNNBlock(
              kernel=_get(weights, f"{pfx}.k"),
              gate_dw_w=_squeeze_conv(_get(weights, f"{pfx}.pregate.net.0.weight")),
              gate_dw_b=_get(weights, f"{pfx}.pregate.net.0.bias"),
              gate_pw_w=_get(weights, f"{pfx}.pregate.net.2.weight"),
              gate_pw_b=_get(weights, f"{pfx}.pregate.net.2.bias"),
              norm_w=_get(weights, f"{pfx}.norm.weight"),
              norm_b=_get(weights, f"{pfx}.norm.bias"),
          )
      
      
      def _build_mlp_block(weights: dict, idx: int, cfg: ReversoConfig) -> MLPBlock:
          pfx = f"layers.{idx}"
          skip_w = _get_optional(weights, f"{pfx}.skip_linear.weight")
          skip_b = _get_optional(weights, f"{pfx}.skip_linear.bias")
          if skip_w is not None:
              skip_w = _T(skip_w)
          return MLPBlock(
              linear_w=_T(_get(weights, f"{pfx}.linear.weight")),
              linear_b=_get(weights, f"{pfx}.linear.bias"),
              final_w=_T(_get(weights, f"{pfx}.linear_final.weight")),
              final_b=_get(weights, f"{pfx}.linear_final.bias"),
              norm_w=_get(weights, f"{pfx}.norm.weight"),
              norm_b=_get(weights, f"{pfx}.norm.bias"),
              skip_w=skip_w,
              skip_b=skip_b,
          )
      
      
      def _build_attn_block(
          weights: dict, idx: int, cfg: ReversoConfig, state_weaving: bool
      ) -> AttentionBlock:
          pfx = f"layers.{idx}"
          ap = f"{pfx}.attention"
          return AttentionBlock(
              q_proj_w=_T(_get(weights, f"{ap}.q_proj.weight")),
              k_proj_w=_T(_get(weights, f"{ap}.k_proj.weight")),
              v_proj_w=_T(_get(weights, f"{ap}.v_proj.weight")),
              o_proj_w=_T(_get(weights, f"{ap}.o_proj.weight")),
              beta_w=_T(_get(weights, f"{ap}.b_proj.weight")),
              q_conv_w=_squeeze_conv(_get(weights, f"{ap}.q_conv1d.weight")),
              k_conv_w=_squeeze_conv(_get(weights, f"{ap}.k_conv1d.weight")),
              v_conv_w=_squeeze_conv(_get(weights, f"{ap}.v_conv1d.weight")),
              o_norm_w=_get(weights, f"{ap}.o_norm.weight"),
              norm_w=_get(weights, f"{pfx}.norm.weight"),
              norm_b=_get(weights, f"{pfx}.norm.bias"),
              n_heads=cfg.n_heads,
              state_weaving=state_weaving,
          )
      
      
      def _build_decoder(weights: dict, cfg: ReversoConfig) -> DecoderHead:
          return DecoderHead(
              head_w=_get(weights, "head.weight"),
              head_b=_get(weights, "head.bias"),
              q_proj_w=_T(_get(weights, "simple_q_proj.weight")),
              q_proj_b=_get(weights, "simple_q_proj.bias"),
              k_proj_w=_T(_get(weights, "key_proj.weight")),
              k_proj_b=_get(weights, "key_proj.bias"),
              v_proj_w=_T(_get(weights, "value_proj.weight")),
              v_proj_b=_get(weights, "value_proj.bias"),
              out_proj_w=_T(_get(weights, "out_proj.weight")),
              out_proj_b=_get(weights, "out_proj.bias"),
          )
      
      
      # ---------------------------------------------------------------------------
      # Weight download
      # ---------------------------------------------------------------------------
      
      _WEIGHT_CACHE: dict[str, dict] = {}
      
      
      def download_weights(url: str, cache_dir: str = "/tmp/reverso") -> dict:
          """Download .npz weights from URL, caching in *cache_dir*."""
          if url in _WEIGHT_CACHE:
              return _WEIGHT_CACHE[url]
      
          os.makedirs(cache_dir, exist_ok=True)
          filename = url.rsplit("/", 1)[-1]
          local_path = os.path.join(cache_dir, filename)
      
          if not os.path.exists(local_path):
              print(f"Downloading weights from {url} ...")
              urllib.request.urlretrieve(url, local_path)
              print(f"Saved to {local_path}")
      
          data = dict(np.load(local_path, allow_pickle=False))
          _WEIGHT_CACHE[url] = data
          return data
      
      
      # ---------------------------------------------------------------------------
      # Autoregressive forecast
      # ---------------------------------------------------------------------------
      
      
      def forecast(
          series: np.ndarray | list[float],
          prediction_length: int,
          weights: dict | str,
          model_size: str = "small",
          config: ReversoConfig | None = None,
          flip_equivariant: bool = False,
      ) -> np.ndarray:
          """Zero-shot time series forecast using Reverso.
      
          :param series: Historical observations (1-D).
          :param prediction_length: Number of future steps to predict.
          :param weights: Either a weight dict or a URL string to .npz.
          :param model_size: One of ``'nano'``, ``'small'``, ``'full'`` (ignored if *config* given).
          :param config: Optional explicit config (overrides *model_size*).
          :param flip_equivariant: Use flip-equivariant averaging.  Helps for
              single-step prediction but can dampen amplitude during multi-step
              autoregressive rollout.  Default ``False``.
          :returns: Array of *prediction_length* forecasted values.
          """
          if config is None:
              config = CONFIGS[model_size]
      
          if isinstance(weights, str):
              weight_data = download_weights(weights)
          else:
              weight_data = weights
      
          model = load_model(weight_data, config)
      
          series = np.asarray(series, dtype=np.float32)
          x_norm, x_min, x_max = preprocess(series, config.seq_len)
      
          forward_fn = model.forward_flip_equivariant if flip_equivariant else model.forward
      
          context = x_norm.copy()
          predictions = []
          remaining = prediction_length
      
          while remaining > 0:
              ctx = context[-config.seq_len :]
              chunk = forward_fn(ctx)
              take = min(config.output_token_len, remaining)
              predictions.append(chunk[:take])
              remaining -= take
              context = np.concatenate([context, chunk])
      
          result_norm = np.concatenate(predictions)[:prediction_length]
          return postprocess(result_norm, x_min, x_max)
      
      
      # ---------------------------------------------------------------------------
      # Warm-up: trigger Numba JIT compilation
      # ---------------------------------------------------------------------------
      
      
      def warmup_jit():
          """Pre-compile the Numba DeltaNet kernel with a tiny dummy input."""
          L, nh, dh = 4, 2, 2
          q = np.zeros((L, nh, dh), dtype=np.float64)
          k = np.zeros((L, nh, dh), dtype=np.float64)
          v = np.zeros((L, nh, dh), dtype=np.float64)
          beta = np.zeros((L, nh), dtype=np.float64)
          _deltanet_recurrence(q, k, v, beta)
      
  • README.md 545 B
    # forecasting-reverso
    
    Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "predict", "time series", "Reverso", or when tabular data with a temporal dimension needs future-value estimation.
    
    See https://github.com/oaustegard/claude-skills/blob/main/.dev-notes/forecasting-reverso-dev-blog.md for dev notes and background
    
  • SKILL.md 5.7 KB
    ---
    name: forecasting-reverso
    description: Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference). Activate when users provide time series data and request forecasts, predictions, or extrapolations. Supports Reverso Small (550K params). Triggers on "forecast", "predict", "time series", "Reverso", or when tabular data with a temporal dimension needs future-value estimation.
    metadata:
      version: 0.1.0
    ---
    
    # Reverso Time Series Forecasting
    
    Produce zero-shot univariate time series forecasts using the Reverso foundation model family (arXiv:2602.17634), implemented in NumPy/Numba for CPU-only container execution.
    
    ## Setup (run once per conversation)
    
    ```bash
    uv pip install numba --system --break-system-packages
    cp /mnt/skills/user/forecasting-reverso/scripts/reverso.py /home/claude/reverso.py
    cp /mnt/skills/user/forecasting-reverso/scripts/load_checkpoint.py /home/claude/load_checkpoint.py
    ```
    
    ## Obtaining Weights
    
    Two paths depending on network access:
    
    ### Path A: Direct download (HuggingFace allow-listed)
    ```python
    import urllib.request, os
    os.makedirs("/tmp/reverso", exist_ok=True)
    url = "https://huggingface.co/shinfxh/reverso/resolve/main/checkpoints/reverso_small/checkpoint.pth"
    urllib.request.urlretrieve(url, "/tmp/reverso/checkpoint.pth")
    ```
    
    ### Path B: User upload (HuggingFace not accessible)
    If the download fails with a network error, tell the user:
    
    > I can't reach HuggingFace from this environment. Please download the checkpoint from
    > https://huggingface.co/shinfxh/reverso/blob/main/checkpoints/reverso_small/checkpoint.pth
    > and upload it here.
    
    Then load from `/mnt/user-data/uploads/checkpoint.pth`.
    
    ### Loading weights
    ```python
    from load_checkpoint import load_checkpoint
    weights = load_checkpoint("/tmp/reverso/checkpoint.pth")  # or upload path
    ```
    
    ## Model Configuration
    
    Reverso Small uses this config (matching the published `args.json`):
    
    ```python
    from reverso import ReversoConfig
    config = ReversoConfig(d_model=64, module_list=["conv", "attn", "conv", "attn"])
    ```
    
    ## Forecasting
    
    ```python
    from reverso import forecast, warmup_jit
    warmup_jit()  # ~2s one-time JIT compilation
    
    result = forecast(
        series=data,               # 1-D array/list of floats
        prediction_length=96,      # how many future steps
        weights=weights,           # dict from load_checkpoint
        config=config,
    )
    ```
    
    The function handles preprocessing (NaN interpolation, padding, min-max normalization) and autoregressive rollout internally.
    
    ### Key parameters
    
    `flip_equivariant=True` — averages forward pass on original and vertically-flipped input. Slightly improves single-step predictions but can dampen amplitude over multi-step rollout. Default is `False`.
    
    ## Input Handling
    
    Accept time series as Python list, NumPy array, CSV column, or inline values. Convert to 1-D float array before calling `forecast()`.
    
    For CSV/DataFrame input, ask the user which column to forecast if ambiguous.
    
    The model's context window is 2048 steps. Series shorter than 2048 are left-padded with the first value. Series longer than 2048 use only the most recent 2048 observations. Provide at least a few hundred real data points for meaningful results — heavily padded context degrades forecast quality because the long convolution kernels process mostly constant input.
    
    ## Visualization
    
    ```python
    import matplotlib.pyplot as plt
    fig, ax = plt.subplots(figsize=(12, 4))
    n = len(history)
    ax.plot(range(n), history, label="Historical", color="#2563eb")
    ax.plot(range(n, n + len(preds)), preds,
            label="Forecast", color="#dc2626", linewidth=2)
    ax.axvline(x=n, color="gray", linestyle="--", alpha=0.4)
    ax.set_xlabel("Time step"); ax.set_ylabel("Value")
    ax.legend(); fig.tight_layout()
    fig.savefig("/mnt/user-data/outputs/forecast.png", dpi=150)
    ```
    
    ## Performance
    
    | Phase | Latency |
    |---|---|
    | numba install (uv) | ~1.6s |
    | Weight loading (.pth) | <1s |
    | JIT warmup | ~2s |
    | Forward pass (L=2048) | ~80ms |
    | 96-step forecast (2 chunks) | ~160ms |
    | 192-step forecast (4 chunks) | ~320ms |
    
    ## Container Environment Limits
    
    Each forward pass takes ~65ms at L=2048. In the ephemeral container, reject batch forecasting requests that would exceed ~1500 forward passes (~100s wall time) to avoid timeouts.
    
    **Detect the container environment** by checking for `/mnt/user-data` or `/mnt/skills`:
    ```python
    import os
    IN_CONTAINER = os.path.exists("/mnt/user-data")
    ```
    
    **Estimate cost before running** when processing multiple series:
    ```python
    n_forwards = n_series * n_windows * max(1, pred_length // 48)
    est_seconds = n_forwards * 0.065
    if IN_CONTAINER and est_seconds > 100:
        # Reject or subsample
        max_series = int(1500 / (n_windows * max(1, pred_length // 48)))
    ```
    
    **Practical limits at ~100s budget:**
    
    | Scenario | Series | Windows | Pred steps | Forwards | Time |
    |---|---|---|---|---|---|
    | Single series, 96-step | 1 | 1 | 2 chunks | 2 | 0.1s |
    | Small dataset (sz_taxi) | 156 | 6 | 48 | 936 | 61s |
    | Medium dataset, short horizon | 300 | 4 | 48 | 1200 | 78s |
    | Large dataset (m4_yearly) | 22974 | 1 | 48 | 22974 | **25min ✗** |
    
    When a request exceeds the budget, inform the user with the estimated time and suggest either subsampling or running locally. For benchmark evaluation of large datasets, recommend running outside the container.
    
    ## Limitations
    
    The model is strongest with periodic or quasi-periodic signals and full 2048-point context. Short series (under ~200 points) are heavily padded and produce degraded forecasts — this is a model limitation, not an implementation bug. Edge cases: binary-valued input (e.g. step functions normalizing to exactly 0/1) and series ending at the exact min-max boundary are out-of-distribution for the training data.
    
    For architecture details, weight mapping, and debugging guidance, read `references/architecture.md`.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related