Tool poisoning
A harmless description can hide instructions that ask the model to read files or credentials.
MCP SERVER TRUST, VERIFIED AT RUNTIME
mcp-sentinel runs third-party MCP servers inside an isolated sandbox, plants fake credentials, and records what happens. Your team receives evidence, not another risk score.
THE SHORT VERSION
A security review that observes behavior, not promises.Static scan + sandbox detonation + capability driftRegistry signal
maintainer unchanged · tool surface changedMost teams still approve them from a registry description. That misses what changes after approval and what the code does when it runs.
A harmless description can hide instructions that ask the model to read files or credentials.
A server approved at one version can quietly gain new tools, schemas, or data access later.
Runtime code can read environment variables and send them away while static metadata looks clean.
Security reports are easy to ignore. A reviewed allowlist PR creates an accountable trust decision.
TrueForge orchestrates. Bright Data supplies change signals. OpenAI reasons over evidence. Qodo reviews the resulting code and trust changes.
The runtime
Runs the agent loop, fans out inspectors, provides Daytona sandboxes, keeps sessions alive, and pauses every external action for human approval.
The registry signal
Scrapes registry pages for allowlisted servers and candidates. It detects version and maintainer changes, but never replaces sandbox evidence.
The reasoning layer
Reads normalized findings and evidence, applies the audit rubric, and proposes one of five constrained verdicts.
The reviewer
Reviews development PRs and the allowlist changes created by the agent. Trust decisions get the same review trail as code.
Registry data decides what deserves attention. The sandbox decides what is true.
THE AUDIT PIPELINE
Scrape registry changes for servers in allowlist.json and new candidates.
Launch one isolated sandbox, plant canaries, list tools, and invoke safe calls.
TRUEFORGE + DAYTONACombine static findings, live behavior, and capability drift into a verdict.
OPENAI + SENTINELPropose an allowlist pull request. Nothing changes until a person approves.
GITHUB + QODOFIVE POSSIBLE VERDICTS
Every verdict carries the exact description, captured request, capability diff, or failure reason that supports it.
THE AUDITOR IS READY
FAQ
An agent that audits third-party MCP servers before your AI agents trust them. It runs each server in an isolated sandbox, plants fake credentials, and watches what the server actually does. You get real evidence instead of a description-based score.
It works in four stages. First it discovers changes from registries. Then it inspects each server in a Daytona sandbox seeded with canary secrets. Then it judges the behavior into one of five verdicts. Finally it opens a pull request against allowlist.json. Trust never changes until a person approves.
TrueForge runs the agent. Bright Data scrapes the registries. OpenAI reasons over the evidence. Daytona provides the isolated sandboxes. mitmproxy captures outbound traffic. GitHub holds the allowlist pull requests. Qodo reviews them.
It scrapes registry pages like npm and Smithery for allowlisted servers and new candidates. It catches version, maintainer, and tool changes. That signal decides which servers need a full sandbox re-audit. Registry data never replaces sandbox evidence.
It is the agent harness. It runs the loop, spins up one inspector per server, provisions the Daytona sandboxes, keeps sessions alive, and pauses every external action for your approval. Sentinel never performs a write on its own.
It reviews the pull requests the agent opens. It checks the code and the allowlist trust changes, so a change to what your agents trust gets the same review as production code. High findings are fixed before merge.
MCP servers receive your credentials and feed their tool descriptions straight into your model as instructions. Approving them from a registry blurb misses tool poisoning, rug pulls that add tools after approval, and secrets that leak only when the code runs. Metadata scanners cannot see any of that.
You get evidence, not another risk score. Every verdict comes with the captured request, the hidden instruction, or the exact capability that changed. The trust decision lives in git where your team can review it. It is behavior observed in a sandbox, not promises read from a description.