"Labeling data should feel like naming stars — not mining coal."
Welcome to SmartLabelBench Nebula, a next-generation auto-annotation workbench that transforms raw, chaotic media into structured, machine-ready datasets. Born from the lineage of prompt-decomposition research, Nebula treats every image, frame, and clip as a puzzle piece waiting to be described — and it uses large language models to deconstruct intent, recompose semantics, and emit pristine labels in one fluid pass.
Whether you are curating a niche dataset of tropical beetles, industrial PCB defects, fashion catalogs, or cinematic drone footage, Nebula scales from a single laptop to a distributed annotation cluster without ever losing coherence.
🌌 Why Nebula Exists
Traditional annotation pipelines ask humans to become machines: click, drag, tag, repeat. Nebula flips the script. It asks the model to become the human — reasoning about your prompt, proposing the label taxonomy, and then applying it consistently across thousands of samples. You remain the curator; the model becomes the tireless scribe.
This repository contains the full workbench: the prompt deconstructor, the label recombination engine, the dataset exporter, and the ancillary tooling needed to operate it responsibly in production.
✨ Feature Constellation
- 🔭 Prompt Deconstruction Engine — Breaks natural-language instructions into atomic semantic units before reassembly.
- 🧩 Recombination Labeler — Merges decomposed fragments into coherent, hierarchical label sets per asset.
- 🖼️ Multi-Modal Ingestion — Accepts images, video frames, PDFs, and remote streams.
- 🌍 Multilingual Label Vocabularies — Native support for 40+ languages, including RTL scripts.
- 📱 Responsive Workbench UI — A tablet-friendly reviewer canvas for spot-checking model output.
- 🛰️ Distributed Worker Mode — Fan out annotation jobs across a fleet of nodes.
- 🗂️ Dataset Exporters — COCO, YOLO, Pascal VOC, JSONL, and a proprietary "Nebula Card" format.
- 🔐 Local-Only Mode — Run entirely offline for sensitive corpora; no telemetry leaves your perimeter.
- 🧪 Active Learning Loop — Model uncertainty feeds back into the human review queue.
- 🕰️ 24/7 Companion Support — A rotating steward team answers questions in the community channel at any hour.
- 🎛️ Plugin Architecture — Register custom LLM backends, custom taxonomies, and custom post-processors.
- 📚 Audit Trails — Every label decision is signed, timestamped, and reversible.
🛰️ Architecture Overview
Nebula is composed of four cooperating layers, each of which can be swapped or extended:
- Ingest Layer — consumes raw assets and normalizes them into a uniform tensor store.
- Reasoning Layer — hosts the prompt deconstructor and the label recombination engine.
- Orchestration Layer — schedules jobs, tracks worker health, enforces quotas.
- Surface Layer — the responsive UI, REST/gRPC endpoints, and dataset exporters.
Each layer communicates through an immutable event bus, meaning you can replay any annotation session from the beginning and observe exactly how a label came to exist. This is essential for regulated industries where provenance matters as much as the label itself.
🚀 Getting Started (Environment Preparation)
Nebula does not assume a specific runtime. Choose the environment that fits your team:
- A containerized deployment for reproducibility.
- A bare-metal deployment for maximum throughput.
- A managed cloud deployment for elastic bursts.
After acquiring the release bundle via the placeholder below, unpack it into your workspace root and consult the bootstrap guide inside the docs/ directory. The workbench ships with a self-diagnostic command that verifies your environment before first run.
No comments yet.