The file tools your agent already has, with inspection in front of them.

Every file an agent opens is content it did not write, arriving already inside the model. Subtext is zero trust file inspection: the same tools your agent calls today, with every byte of file content run through the engine before any of it is returned.

block  review  pass 0  cloud calls 0  third-party packages
Where it sits

Not a policy the model is asked to follow.

The enforcement is not a system prompt, a rule in the agent’s instructions, or a classifier the model consults. It is the tool itself. When the agent calls for a file, the bytes are read, inspected, and either returned or withheld. The model never gets a chance to decide.

Client Your agent
→
Inspection Subtext reads every byte first
→
Source Every file it is asked to open

The read and metadata tools carry the same names your agent already calls, so they go in as a replacement for the uninspected ones rather than as an addition. One entry in your MCP config, restart the client, done. Python standard library only: no package install and no build step. It is a client of the engine rather than a copy of it, so Subtext runs alongside it and has to be reachable.

One entry, and the tools are inspected"subtext-fs": { "command": "python", "args": ["/path/to/subtext_fs.py"], "env": { "SUBTEXT_URL": "http://localhost:8000/v1/score", "SUBTEXT_API_KEY": "your key", "SUBTEXT_MODE": "enforce", "SUBTEXT_FAIL": "closed" } }
Client support

Validated on Claude Code, OpenAI Codex, xAI Grok, GitHub Copilot, Microsoft VS Code Copilot and Google Antigravity, each on the same connector with no changes between them. Any client that speaks MCP over standard input and output should work, but those six are what we have tested. Tell us which one you run and we will tell you honestly whether it is on the list.

For servers you cannot replace

The same inspection core also runs as a proxy that sits in front of an MCP server you do not own, parsing the protocol frames and applying the same verdicts to whatever that server returns. Ask us about it if the connector you need to gate is somebody else’s.

What it gates

Four questions, asked of every byte before it returns.

Each one is an operator dial, not a fixed opinion. The policy lives in the engine, so a change applies to every consumer of it at once, and the tools honor that decision rather than second-guessing it. You set what each verdict does here, including whether a review verdict returns the file or refuses the read.

Prompt injection Instructions wearing the shape of data A file is untrusted content, and the model reads all of it. Subtext reads it first: the visible text, the concealed text, and the text inside whatever the file is carrying.
Sensitive data What you did not mean to send Reading a file is how its contents enter the context window, and from there a request to whoever runs the model. Identifiers, payment data, credentials, and your own listed terms are caught in the file before the read returns, rather than after the context is built.
File type Types you never agreed to accept Set which classes of content an agent is permitted to open at all. Type is resolved from the bytes, not from the extension or the header the source declared, so a mislabeled file is a finding rather than a free pass.
Language Content your deployment does not read Declare the languages in scope, so content your deployment does not read is not judged by checks that were never calibrated to read it.

All four run against the same extracted content. A rule about file types is worth little if the file was never opened.

Injection is the one this boundary creates. The other three are policy you would want anywhere files move, enforced here as well.

Coverage

Reads and fetches.

Content has no direction. A payload arriving and a secret leaving are the same inspection. Here is exactly which tools inspect content and which do not, stated plainly so a demo does not imply coverage it does not have. The connector reads; it has no write tool.

ReadAny file the agent opens
The bytes are read and scored before anything is returned. Media comes back as base64 with its type, so an image or an audio file is opened and inspected rather than handed over as an opaque blob.
On by default
FetchContent pulled from a URL
This one is an addition rather than a replacement, because the filesystem server it stands in for has no fetch tool. It widens what the agent can reach, so read the egress note below before you turn it loose. A URL is a file with extra steps: whatever comes back is downloaded to a size ceiling, scored, and returned only if it passes. That means the page’s own HTML as much as anything pulled down through it, and an attachment or a download gets exactly what a file off the disk gets. The bytes are what gets typed. The URL’s path and the type the server declared are recorded as claims, not trusted as answers.
On by default
MetadataNames, sizes, and structure
These return filenames and stat information. There is no file content in the response, so there is nothing to inspect and none is claimed.
No content
Why content, not text

Reading the file is the whole job.

Gateways in this space read the text of a response, and some also have a document path. As of August 2026, we have not found one that documents checking the declared type against the actual bytes, or failing closed on a format it cannot parse. Those two gaps are where the following live. Hover a card.

Declared · an ordinary document A memo, a report, a resume.
Actual Instructions aimed at the model reading it, sitting in the body or concealed where a person would not look. The agent was told to read this file. It will read all of it.
Hover →
Declared · the extension, or a Content-Type header The source says what this is.
Actual Both are supplied by the thing being inspected. Route on either and a file can pick its own level of scrutiny. Subtext resolves type from content and treats a mismatch as a finding in its own right.
Hover →
Declared · a link the agent was given Fetch this and summarize it.
Actual Whatever the server on the other end decides to return, entering context under the authority of a URL somebody pasted. The markup itself is content, including text styled or positioned so a reader never sees it and the model reads it anyway. Attachments and downloads get the same inspection a file off the disk gets.
Hover →
Declared · a photo Somebody’s ID badge, photographed.
Actual The number on it is text, at whatever angle the badge was held. A multimodal model reads it either way. Decoded and read here first, it comes back as a sensitive-pattern match, where anything iterating text blocks sees an opaque image and forwards it. Needs the image path, which is a deployment-profile choice.
Hover →

Same engine as the file and mail boundaries. The transport changed. The question did not.

How it enforces

Blocking without breaking the agent.

Failing silently leaves the agent guessing. Rewriting the file destroys the evidence. Subtext returns something the model can read, marked as a tool error so the agent treats it as a failed read rather than as data, and names the file it refused.

What the model receives instead[Subtext block: prompt injection. Content withheld from the model. path=quarterly_memo.txt]
Fail closed Uninspected is not the same as clean If the engine is unreachable, the read is refused. Fail-open is available for availability-sensitive deployments, as an explicit setting, and the reason is always in the audit record. It is never the silent default.
Watch first Monitor mode before enforce mode Monitor returns everything and records the verdict, so you can see what your agents have actually been reading before you turn enforcement on. The first day of logs is usually the interesting part.
Evidence What the content was, not just that a tool ran One record per inspection decision: path, direction, verdict, the policy action applied, score, and reason. It goes to the client’s own logs, so it lands wherever you already collect them. A behavioral layer can tell you a tool was called. This tells you what came back.
The component itself

Nothing listening, and nothing to configure.

Adding a security component should not add attack surface, and it should not add a second thing to keep in sync. We spend a good deal of time reporting bugs in other people’s MCP integrations, and the recurring ones are network-facing: a service bound wider than intended, a missing authentication check, an origin nobody validated. This opens no port, and it needs no list of approved paths to keep current.

Local A child process of your client It speaks over standard input and output to the one client that launched it. There is no listener to authenticate, no origin to validate, and nothing on the network that can reach the connector.
Unscoped by default Inspection, not access control By default it narrows nothing: it changes what comes back, not what the agent can reach. Pass directories on the command line and it refuses anything outside them, but that is an option you turn on rather than the mechanism. An allowlist is a thing to maintain and a thing to get wrong, and the file that hurts you is usually the one nobody thought to put on the list.
Small One file, standard library No third-party packages, so there is no dependency tree to audit. You can read the whole thing in one sitting, which is a reasonable thing to expect of something you put in this position.
Limits

Where this stops.

Enforcement is only as complete as the tools it covers, and inspection is only as complete as what the content can show. Worth stating before you deploy it, not after.

Prose with no marker on it Instructions hidden in content have shapes: a forged system boundary, a forged interface element, a payload encoded so the reader never sees what the model reads. Those are content, content is inspectable, and inspecting them is the job we take. What we do not claim is the other half. A plausible complaint, a helpful-sounding suggestion, an ordinary sentence that happens to redirect an agent, with no forged role and no imperative in it. There is nothing structurally wrong with that text, and a deterministic layer that tried to catch it would be guessing at intent. That half belongs to the model. A model that acts on instructions it read in content it was handed is not ready for enterprise use, and no gateway fixes that from the outside.
Uninspected file tools alongside these This is the one that catches people. If the tools you are replacing are still registered next to these, the agent has an unsupervised path to the same files and will use it. Replace them, do not add to them.
Other tools that reach the same content A built-in web fetcher, a browser tool, and a built-in shell are implemented by the client, not by a connected tool, so the client runs them and puts the result straight into context. That content is not inspected. The connector ships inspected replacements for the fetcher and the shell, so what comes back through them is scored before the model sees it. For the shell that means the output, not the command, so a command that writes has already written by the time anything is checked. Turn the native tools off so the inspected ones are the only route. The browser tool has no replacement, so inspect that traffic where it crosses a boundary you control. The same engine runs at the network gateway.
Other connectors and hosted integrations A ticketing or mail connector returning an attachment is a different channel. Integrations authorized inside a vendor’s own settings run on their infrastructure, and nothing local is in that path. Those need the proxy shape or a different enforcement point.
Outbound fetches need an egress policy URL fetching is held to http and https on every hop, redirects included, and capped by size, but nothing constrains where it will reach. On a workstation that is usually fine. Anywhere with internal services or a cloud metadata endpoint, put an egress policy in front of it, the same as you would for any other component that makes outbound requests.
Content that needs credentials URL fetching sends no credentials, so it cannot retrieve anything behind a login. Content that requires authentication has to reach the agent some other way, and that way is not covered here.
Read further

The work behind it.

Swap the tools and watch.

Point it at a directory your agents already read, leave it in monitor mode, and let it record. It returns everything and logs a verdict for each file. Then decide what you want blocked.

Runs on your side of the wire. No cloud account, no feed to keep current.