Inference Aiops
Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
- Transport
- Not stated
- Package
- —
- Registry id
- io.github.AIops-tools/inference-aiops
No install snippet on purpose. A working MCP config is a command, its arguments and an environment block — the last two are where API keys live, so this catalogue never stores them and cannot publish them. Follow the link above for the authors' own instructions.
Disclaimer: Community-maintained open-source project. Not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor. Product and trademark names belong to their owners. MIT licensed.
Governed AI-ops for GPU inference clusters — vLLM (OpenAI API + Prometheus
/metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process
serving engines SGLang and TGI (Text Generation Inference) — with a
built-in governance harness: unified audit log, policy engine, token/runaway
budget guard, undo-token recording, and descriptive risk-tier labels on every
audit row. It parses each engine's Prometheus /metrics directly (no Prometheus
server required) and
probes the Ray dashboard independently. A bearer token is optional (many
stacks run open).
Serving engines. vLLM is the flagship (full Ray Serve control plane: scale, drain, autoscale, LoRA, hot-swap). SGLang and TGI are supported for engine-agnostic observability — health, running-model identity, request-latency metrics, queue depth, and latency RCA — read from each engine's own endpoints and metric names. Being single-process servers, they have no Ray-shaped scale/drain API: those writes return a teaching error pointing you at a real horizontal-scale layer (Ray Serve / Kubernetes / a load balancer).
What it does
The flagship value is root-cause analysis, wrapped in guarded reads and writes:
diagnose_latency_spike(flagship RCA) — when TTFT/TPOT/e2e latency climbs, it correlates queue depth (running vs waiting), KV-cache pressure / preemptions, and prefix-cache locality into a ranked cause plus the specific knob to turn (add replicas, raisemax-num-seqs, fix routing, enlarge KV cache). Every flag is a number, not a black-box verdict.diagnose_low_utilization— the inverse: idle GPUs, over-provisioned replicas, or routing that strands a cache-warm replica → what to scale down.- Prometheus-native — reads vLLM's
/metricsendpoint directly; no Prometheus/Grafana deployment needed. - Governance-grade — the first governance-grade entrant in this niche: audit + budget + risk-tier approval + undo-token + prompt-injection sanitize, with dry-run + double-confirm on the fragile prod ops (scale-down, scale-to-zero, drain, redeploy, hot-swap) the community reports as dangerous.
- Laptop self-test — ~80% of the tool self-tests free: vLLM on a single GPU
or CPU-mock + Ray in one local container (
ray start --head).
What this tool does, and does not, decide
It delivers inference-cluster operations — reads and writes — accurately and efficiently, and records every one of them. It does not decide whether a write is allowed to happen. That is the agent's judgement, or the permission of the environment you connect it with: restrict the network path so it can only reach the read/metrics endpoints, or run the Ray dashboard without its job-submission API, and the writes fail at the server — the place that actually owns the permission.
So there is no read-only switch, no policy file, no approval gate to configure.
The one thing the tool guarantees is that nothing is silent: every call, over
MCP and over the CLI alike, lands an audit row in
~/.inference-aiops/audit.db, and destructive writes still capture their
before-state and record an inverse where one exists.
Each tool declares a
risk_level, kept in agreement with its[READ]/[WRITE]documentation tag by a test, and carried into the audit row as a descriptive tier — so a reviewer can see at a glance that a row was a high-risk scale-to-zero. It is a label, not a gate.
Running a smaller / local model? See agent-guardrails.md — it lists the guardrails this tool enforces for you (so you don't spend prompt budget restating them) and gives a ready-made system prompt for what's left.
From the project's README.