BIP NYC

collapse
Home / Daily News Analysis / Microsoft tells court Copilot rarely copies books while selling OpenAI's most capable model

Microsoft tells court Copilot rarely copies books while selling OpenAI's most capable model

Sep 05, 2026  Twila Rosenbaum  3 views
Microsoft tells court Copilot rarely copies books while selling OpenAI's most capable model

Microsoft is pressing two starkly different messages about its AI business this week. In a New York federal court, it argued that its AI-powered Copilot assistant does not meaningfully reproduce copyrighted books and that training large language models on books is fair use. Just two days before that filing, however, the company started selling what it calls its most capable model ever: GPT-6 Astra, an AI system from OpenAI that can read screens, act inside software interfaces, and meets OpenAI's most severe cybersecurity classification.

The coincidence of timing has turned the twin announcements into a study in contrasts. In the legal arena, Microsoft emphasizes the limits of its AI systems, saying an expert found just 24 matching responses out of 8.2 million Copilot conversations. In the commercial arena, Microsoft is selling a model that it places at the very frontier of what OpenAI has built, with capabilities that include autonomous computer use and a designation that raises systemic risk concerns.

The case is part of a broader multidistrict copyright litigation before Judge Sidney Stein in the U.S. District Court for the Southern District of New York. The litigation consolidates claims from authors and publishers who say Microsoft and OpenAI unlawfully used their books to train AI models. The claims involve 212 books, though Microsoft says the expert found nothing at all for 202 of them. The company is now asking the court to decide the entire question in its favor as a matter of law, a motion for summary judgment that would avoid a trial on the core fair use question.

Microsoft's memorandum argues that the few matching passages, which it calculates as 0.00029 percent of all Copilot conversations examined, are de minimis and cannot support a finding of copyright infringement. It further argues that the training process itself is fair use under long-established U.S. copyright principles, which allow the use of copyrighted materials for transformative purposes. In Microsoft's view, the creation of large language models resembles the kind of intermediate copying that courts have permitted when the final output does not supersede the original work.

A Rare Admission of Matching Outputs

The figures presented in the motion offer a rare glimpse into how often a commercial AI assistant actually generates text that matches passages from books used in training. According to Microsoft, an expert engaged for the litigation searched for passages from the 212 books at issue across 8.2 million Copilot interactions. The search found 24 matching responses total, a number Microsoft says is far too small to cause any market harm or to justify a finding that Copilot is a substitute for the books.

For the vast majority of the books in the case, Microsoft says there were no matches at all. That point is central to its summary judgment argument. The company says the plaintiffs cannot show that their specific works were reproduced by the model in any meaningful way, and therefore their claims should collapse. The motion also contends that even if some outputs are similar to copyrighted text, those outputs are not the kind of direct or substantial copying that copyright law is designed to prevent.

The plaintiffs, which include prominent authors and publishers, have taken a different view. They argue that the use of copyrighted books to train AI models is itself an act of infringement, regardless of whether the model later reproduces the text. They also claim that the AI companies have unjustly profited from millions of copyrighted works without permission or compensation. In earlier rulings, Judge Stein has allowed the copyright claims to proceed past the pleading stage, rejecting arguments that the use of books in training is categorically protected by fair use.

Microsoft Launches GPT-6 Astra

While Microsoft was preparing its defense in New York, it announced on September 3 that GPT-6 Astra, a new model from OpenAI, was available through the Foundry Limited Access Program. The model is priced at $10 per million input tokens, positioning it as a premium offering for enterprise customers. Microsoft describes Astra as its most capable model to date, and the company's Azure blog highlights Astra's ability to perform computer use tasks.

The model can read on-screen information from a wide variety of applications, act inside approved interfaces, update records, test software, and navigate development tools. That kind of capability represents a significant step beyond text generation and into the realm of autonomous digital labor. Enterprises could conceivably use Astra to automate repetitive tasks such as data entry, quality assurance testing, and software navigation, but the model also requires careful safeguards to prevent mistakes or misuse.

Microsoft says it has built a range of protections around the model, including scoped credentials, human checkpoints for consequential actions, and activity records. The company also states that prompts and outputs are not used to train the underlying models. Those measures are intended to reassure enterprise customers who may be wary of granting an AI agent access to their internal systems and data.

A Critical Cybersecurity Designation

Perhaps the most striking detail in Microsoft's announcement is the security classification attached to GPT-6 Astra by OpenAI. Under OpenAI's Preparedness Framework, the model has been rated as meeting the Critical cybersecurity threshold, the first time OpenAI has assigned that designation to a model. The Critical level means the model could be used to find unknown security flaws and develop exploits without step-by-step human guidance. In practical terms, that suggests the model has a sophisticated understanding of software vulnerabilities and could assist both defenders and attackers.

OpenAI's Preparedness Framework is an internal evaluation system that scores models on several risk categories, including cybersecurity, biological threats, chemical threats, and societal influence. A Critical rating does not mean OpenAI has observed malicious behavior; it means the model's capabilities are high enough that the company believes special precautions are required. For Microsoft, offering such a model through its cloud platform raises questions about how enterprises will deploy the technology and whether they can monitor its activity sufficiently.

The juxtaposition of the two announcements is notable. In the copyright lawsuit, Microsoft is telling the court that Copilot's outputs are so limited that the model is nothing like a copying machine. In the product launch, Microsoft is telling the market that GPT-6 Astra is capable of complex, autonomous actions with potentially significant consequences. Both statements can be true, but they point in opposite directions when regulators and judges ask how powerful these models have become.

European Law Looks at Capabilities, Not Just Outputs

The legal tension is even more pronounced across the Atlantic. In Europe, the debate over AI and copyright is not centered on whether a model occasionally reproduces training text. Instead, European courts are beginning to examine whether the model itself contains copies of protected works in its stored parameters, or weights, and whether that storage alone infringes the reproduction right.

On July 31, the Munich Regional Court ruled that works memorized in a model's stored parameters infringe the reproduction right at that point. The case, GEMA against Suno, is under appeal, but the ruling has already sent a signal through the European AI industry. If a copy of a musical work exists in the weights of a generative model, the court reasoned, it does not matter how rarely that copy surfaces in the model's output. The unauthorized reproduction occurs when the model is trained and the parameters are saved.

This approach contrasts sharply with the U.S. fair use doctrine, which considers factors such as the purpose of the use, the nature of the copyrighted work, the amount used, and the effect on the market. Where a U.S. court might focus on the small number of matching outputs, a European court might focus on the fact that the model has internalized and stored some expression of the works during training. The two legal systems are built on different assumptions about what is harmful and what is permissible.

Microsoft's summary judgment filing would face a much different reception in Europe. The company's argument that 24 matches in 8.2 million conversations is negligible would not necessarily defeat a reproduction right claim under the Munich ruling, because the alleged infringement is not the later output but the earlier copying in the training process. Even if no verbatim output ever occurred, the storage of expressive elements in the model's parameters could still be considered an infringement.

The AI Act's Systemic Risk Duties

The launch of GPT-6 Astra also raises questions about how the European Union's AI Act applies to models with Critical cybersecurity capabilities. The AI Act, which entered into force in stages, attaches special obligations to general-purpose models that carry systemic risk. Those obligations are defined in terms of capability rather than output frequency.

Under Article 55 of the AI Act, providers of general-purpose models classified as having systemic risk must conduct documented adversarial testing. They must report serious incidents to the AI Office without undue delay. They must also secure the model and its physical infrastructure. None of these duties depend on how often the model repeats a passage from a book. A model that never reproduces a copyrighted sentence could still be subject to the full range of systemic risk obligations if its capabilities pose a serious risk to public safety, public health, or cybersecurity.

That distinction is crucial for understanding why Microsoft's simultaneous legal and commercial messages can create confusion. The legal message is designed to minimize the model's copying behavior, implying that these systems are not sophisticated enough to infringe meaningfully. The commercial message is designed to maximize the model's perceived capabilities, implying that these systems can act as autonomous agents with advanced reasoning and computer skills. Both cannot easily fit into a single regulatory narrative, especially in Europe where capability triggers obligations and memorization triggers infringement.

An EU Data Zone Gap

Microsoft's launch of GPT-6 Astra also includes a notable geographic omission. Astra is initially available with Global and US data zone deployments for the Foundry Limited Access Program. Microsoft has offered an EU data zone on Azure since November 2024, meaning customers in the European Union can keep their data within the EU region to comply with local data protection rules. No such option was announced for GPT-6 Astra, despite the model being sold in a global market that includes European enterprises.

For European customers, the absence of an EU data zone for Astra may be a serious concern. Many companies in the EU are already wary of sending data outside their own region because of GDPR restrictions, national security concerns, and growing political pressure to reduce dependence on American cloud providers. The lack of an Astra deployment in the EU data zone could slow adoption among European enterprises that would otherwise be interested in the model's capabilities.

Microsoft has not explained why the EU data zone was not included at launch. It is possible that the company is waiting for regulatory approvals or that the infrastructure for hosting such a powerful model in the EU is still being prepared. But the timing is awkward, as European policymakers debate the role of foreign AI models in essential sectors and call for stronger digital sovereignty. A company that is simultaneously defending its AI models in court and selling its most advanced model to the world is now being observed by judges and regulators who are asking entirely different questions about how AI should be governed.

The coming months will show whether the fair use argument succeeds in New York and whether the Munich ruling survives appeal. Those outcomes could shape the boundaries of AI training for years. In the meantime, Microsoft's week offers a vivid example of the contradictory pressures facing major AI companies: the need to reassure courts that their models are harmless, while also convincing customers that their models are powerful enough to transform the way businesses operate.


Source: TNW | Artificial-intelligence News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy