
Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data"
Quick Answer
Microsoft's MAI models were trained on unlicensed web data, contradicting its claims of using only 'clean and commercially licensed data.' This practice mirrors that of other AI companies, relying on fair use while placing the onus on website owners to block crawlers.
Key Points
- Microsoft's MAI models utilize unlicensed data sources like Common Crawl.
- The company claims to provide 'enterprise grade' data but does not adhere to this.
- Microsoft's approach shifts the responsibility to site owners to block data crawlers.
- This practice is common among AI labs, raising ethical concerns.
📖 Reader Mode
~1 min readMicrosoft partly trained its new MAI models on unlicensed web data. The technical paper shows Microsoft used Common Crawl, among other sources, as Simon Willison noted. Microsoft had previously claimed the MAI models were trained only on "enterprise grade, clean and commercially licensed data."

Like other AI companies scraping the web, Microsoft is likely relying on fair use. The paper describes the data as a "mixture of publicly available and licensed human-generated data." For web data, Microsoft says it uses "a proprietary crawler that respects the Robots Exclusion Protocol (robots.txt) and related meta-tag and HTML controls, enabling site owners to manage how content on their sites is accessed and used."
That puts the burden of protecting content on site owners, like assuming anyone who doesn't lock their door consents to a break-in. Fair use remains contested, and courts are still sorting it out. In short, Microsoft does what every other AI company does, yet sells its training data as especially "clean." It isn't.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

