Self-hosted · works with Claude, Gemini and more

Your teams already use AI. Your sensitive data shouldn't be part of it.

TokenVeil strips names, IPs, secrets and identifiers before a message reaches an external model, then restores them in the response. Your teams keep their habits. The real data never leaves your infrastructure.

See how it works

Deployed in-house. No data is sent to a third party for the anonymization step itself.

Preview sent to the AI
>client=Sophie Marchand
>email=sophie.marchand@client.com
>server=10.42.8.17
>apikey=sk-live-4f8a9c2e1b...

Works with the models your teams already use

Claude (Anthropic)Gemini (Google)OpenAIMistralGitHub ModelsAzure OpenAIVertex AIAmazon Bedrock

The risk isn't AI. It's what you send it.

Logs pasted into a chat contain more than expected

An internal IP address, a client name, an API token buried in a stack trace: it all goes out with the rest of the message, without anyone reading it line by line.

Consumer AI providers weren't built for this

Terms of use, data retention, different jurisdictions: what's acceptable for personal use isn't necessarily acceptable for a company's data.

Banning AI doesn't work

Teams find a way to use it anyway, often on a personal account, outside any oversight. The real issue is what goes into the message.

Three steps, zero change in habits

Anonymization happens between the user's keyboard and the model. No one needs to learn a new tool.

01

Connect your AI accounts

Claude, Gemini, OpenAI, Mistral, or another supported provider: each user links the account they already use. Once, per user.

02

Your teams write as usual

No change in how they work. The text is anonymized before leaving your infrastructure, in a few milliseconds.

03

The response comes back complete

The real data is restored in the displayed response. The external model never saw it.

Built for teams that already answer to someone

Detection

Detection that goes beyond regex

Microsoft Presidio and spaCy NER (French and English), plus custom rules for IPs, IBANs, cards, and API keys. CamelCase identifiers get split and checked, Apache and Nginx log lines are recognized, and in JSON logs a sensitive key (name, phone, iban) triggers tokenization whatever the value looks like.

Custom business keywords

Internal project names, codenames, identifiers specific to your organization: add your own rules, beyond generic detection.

Leak rate measured continuously

0% leak rate, measured across more than 3,300 sensitive values with a public, reproducible method: an annotated test set plus random-seed fuzzing. Not a frozen benchmark: an independent verifier and a CI gate set at zero leaks re-check the number every time the engine changes.

Documents

Documents anonymized too

Attach a Markdown, Word, Excel, or PDF file, including a scanned PDF (OCR built in). The content is extracted, anonymized, and sent through the same pipeline as a typed message. The AI reads the document, never the real names, emails, or numbers inside it.

Hand off a clean copy

Export a redacted version of a Word or Excel file, same format, same structure, sensitive content replaced by tokens. Hidden metadata and embedded thumbnails are stripped too, so a redacted document doesn't leak through its preview image.

Governance

Audit log, not the data

Every anonymized message is tracked (who, when, how much data per category) without ever storing the real value in the log itself.

Turn off what you don't need

Some companies don't want internal IPs touched at all. Each detection category can be switched off for the whole deployment from the admin panel, no code change involved.

LDAP with multi-tenant seat allocation

Local accounts or your company directory. Several LDAP groups can share one instance, each with its own seat limit, separate from the global license cap.

Deployment

Eight AI providers, no vendor lock-in

Claude (directly, or billed through Vertex AI or Bedrock), Gemini, OpenAI, Mistral, GitHub Models, Azure OpenAI. Switch mid-conversation, the anonymized history follows.

Self-hosted, infrastructure you control

The anonymization engine runs on your side. No sensitive data passes through a third-party service before being processed.

Anonymize here, work anywhere

Anonymize a text, paste it into your own tool (Claude Code, ChatGPT), paste the reply back to translate it. The mapping stays encrypted on the server: only an expiring, revocable, audited share code ever circulates.

The tool, in real conditions

Screenshot of TokenVeil anonymizing a message in real time before sending it to Claude or Gemini

Chat view, live preview of the anonymization as you type.

Built to pass a security review

TokenVeil is designed so a CISO can answer the question "where does our data go" without having to take an external provider's word for it.

  • Hosted on your side: no sensitive data passes through a third-party service before processing
  • Token-to-real-value mapping encrypted at rest, per session
  • Audit log kept separate from the data: traceability without risk of leaking it into the log itself
  • LDAP / Active Directory support, including multiple groups on one instance with separate seat limits
  • Detection categories can be turned off per deployment, so you control exactly what gets touched
  • Redacted documents have their hidden metadata and embedded thumbnails stripped too, not just the visible text
  • Source code available for review: no black box in how data is handled

Free and open source

The core engine is public on GitHub: audit the code and evaluate the tool with the free Community edition, which covers deterministic data (IPs, emails, secrets). Enterprise adds ML detection that catches people and client names in free text: that's the edition measured at 0% leakage.

View the code on GitHub
Community and Enterprise editions compared
CapabilityCommunity (regex, free)Enterprise (Presidio + spaCy NER + ML)
Deterministic data (IPs, emails, secrets, IBANs, cards, French NIR, phone numbers)IncludedIncluded
Names after a civility title (Mr. Dupont)IncludedIncluded
People, client, organization, and place names in free textNot includedIncluded
CamelCase, User-Agent, and query-param heuristicsNot includedIncluded
Measured leak rate~0% on deterministic categories0% across 3,340+ values, prose names included
Multi-provider chat, local accounts, admin, audit logIncludedIncluded
Documents (Word, Excel, PDF, OCR) and redacted copiesNot includedIncluded
LDAP / Active Directory, multi-tenant seatsNot includedIncluded

Try TokenVeil in real conditions

Leave your email, we'll reach out directly to talk about your use case. No generic pricing shown here, we discuss it together.

contact@tokenveil.eu

Share this with your team: download the one-pager (PDF)