MistralAI-Training

Mistral AI · Training claims robots compliance vendor-doc verified July 2026

What it does

MistralAI-Training is Mistral AI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Crawls web content to build datasets for training Mistral generative models. Full UA: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots). THE token a training-opt-out robots.txt needs for Mistral — MistralAI-User/-Index do not cover training. No published IP list means UA-spoof verification is impossible; new token since our 2026-07-19 snapshot.

How to verify a hit is really MistralAI-Training

Verification method (July 2026): none published — no IP list (mistral.ai/mistralai-training-ips.json 404s, checked 2026-08-22), unlike the vendor's other two tokens

No published verification path. Mistral AI publishes no IP list, rDNS pattern, or crawler-source documentation for MistralAI-Training as of our July 2026 review — a UA hit cannot be positively confirmed as Mistral AI infrastructure. Treat every MistralAI-Training log line as unverified: UA strings are freely spoofed, and for this bot there is nothing to check them against. General triage still applies — see verify by IP.

Allow or block in robots.txt

Match the exact token MistralAI-Training in robots.txt:

# Block MistralAI-Training site-wide
User-agent: MistralAI-Training
Disallow: /
# Explicitly allow MistralAI-Training
User-agent: MistralAI-Training
Allow: /
Compliance is claimed, not independently confirmed. Mistral AI states this bot respects robots.txt; treat the directive as the first line of defense and watch your logs after deploying it.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether MistralAI-Training may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.

Vendor documentation: https://docs.mistral.ai/robots/

The Mistral AI family

Mistral AI operates 3 distinct tokens in this roster, and they do not block each other: a robots.txt line for MistralAI-Training does nothing to its siblings — each needs its own User-agent: entry (the robots.txt guide has combined patterns).

TokenRolerobots.txt
MistralAI-UserUser actionclaims robots compliance
MistralAI-IndexAI searchclaims robots compliance

Frequently asked questions

What is MistralAI-Training?

MistralAI-Training is Mistral AI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does MistralAI-Training respect robots.txt?

Mistral AI claims MistralAI-Training respects robots.txt, but compliance is not independently confirmed.

How do I block MistralAI-Training?

Add 'User-agent: MistralAI-Training' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

How do I verify MistralAI-Training traffic by IP?

You can't with confidence — Mistral AI publishes no IP list or rDNS pattern for MistralAI-Training as of our July 2026 review, so UA-only hits stay unverified.

Related bots

Same vendor, then other training crawlers — the full directory profiles all 34.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (25/34 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.