- Claude Opus 4.6 introduces a context window of up to 1 million tokens in beta and reinforces its role in scheduling and long agentic tasks.
- The model leads benchmarks such as Terminal-Bench 2.0, Humanity's Last Exam, GDPval-AA and BrowseComp, outperforming GPT-5.2 and Gemini 3 Pro in several specific tests.
- It incorporates features such as adaptive thinking, effort levels, context compaction, agent teams in Claude Code, and improvements in Excel and PowerPoint.
- It maintains a price of $5/$25 per million tokens (entry/exit), is available via web, API and major clouds, and strengthens security and cybersecurity controls.
The arrival of Claude Opus 4.6 It marks a new step in the race for high-end artificial intelligence models. Anthropic has taken the previous version as a base and taken it a little further. in the areas that companies demand most: Reliable programming, long-term tasks, and handling of large volumes of information without the system losing its thread halfway through.
More than just a chatbot, Opus 4.6 wants to position itself as a work collaborator capable of working for hours on the same problem, reviewing its own code, exploring entire repositories, or analyzing lengthy documents. All while maintaining a consistent pricing structure and without major surprises for those already working with Anthropic models through the large clouds.
What is Claude Opus 4.6 and why does it matter?
The launch of Claude Opus 4.6 It is presented as the direct evolution of Opus 4.5, with one particularly striking new feature: a context window of up to one million tokens in beta phaseThis allows you to ingest and reason about conversations, code repositories, or reports that previously had to be broken down into much smaller fragments, reducing the risk of losing important instructions along the way.
Anthropic defines this model as its most advanced system to dateIt's geared towards teams that need to go beyond a simple, one-off response. The idea is that it can handle long-term projects —from initial research to implementation of solutions— acting as an agent that maintains consistency even when the material to be processed is the size of several novels or an entire repository.
In the competitive arena, the company places it in the same league as OpenAI's GPT-5.2 o Google Gemini 3 ProAnd it claims that in several technical tests, the new model outperforms its rivals. Although benchmarks don't always perfectly reflect real-world use, they do give an idea of the types of tasks in which Opus 4.6 aims to excel.
For European companies, professional firms, or technology startups, this approach fits a very specific need: automate complex knowledge work without having to multiply staff, especially in contexts of tight costs and increasingly shorter delivery times.
Key capabilities: from coding to office work
One of the pillars of the new model is its performance in programming and code debuggingOpus 4.6 is designed to navigate large codebases, understand internal dependencies, review changes, and detect errors that previous versions overlookedAnthropic insists that the model is capable not only of generating code, but also of reviewing its own work and correcting it.
Beyond pure development, the model has been trained to handle Financial analysis, advanced research, and working with documents, spreadsheets, and presentationsThis includes the ability to combine several tasks in a chain: for example, extracting data from reports, processing it in Excel, generating conclusions, and presenting the result in a coherent presentation.
In the day-to-day operations of a company, that kind of continuous flow is what separates a "curious" AI assistant from a tool that truly saves hours of repetitive workAnthropic's stated goal is for Opus 4.6 to be able to handle entire work blocks, not just small, isolated queries.
The model is also designed to work within multitasking environments where several requests are launched sequentially, as happens in collaboration platforms or corporate productivity suites. In these contexts, the ability to retain a memory of what has been done previously is critical to avoid having to repeat explanations over and over again.
1 million token context window and long-term memory
The most talked-about technical change of Claude Opus 4.6 is his context window of up to one million tokens on the developer platform. In practical terms, this means you can work with large code files, contracts, technical documentation, or chat logs without having to split them into artificial blocks, something that was previously a common headache.
This leap in capacity seeks to address a fundamental problem: the so-called “context rot”As a session drags on, many models begin to contradict themselves, forget previous decisions, or stop following initial instructions. According to published data, Opus 4.6 significantly improves in tests where specific information must be found hidden within hundreds of thousands of words, with success rates much higher than previous generations.
The key is not just being able to "store" more text, but use it effectivelyAnthropic argues that the model maintains logical consistency in very long interactions, something relevant when working with audits, due diligence, extensive legal reviews or system migrations, common tasks in European consultancies and corporate departments.
For particularly long sessions, the system incorporates a mechanism of context compactionWhen the conversation approaches a defined threshold, the model automatically summarizes the oldest parts and replaces them with condensed versions, freeing up space to continue working without losing the key elements that were agreed upon at the beginning.
Adaptive thinking and effort levels: more control for developers

Along with the model, Anthropic has introduced new options in the Claude Developer Platform to adjust how AI reasons. The so-called adaptive thinking It allows the system to decide how much depth of reasoning to apply depending on the task, instead of operating with a simple "extended mode yes or no" switch.
In addition, the following have been defined: four levels of effort —low, medium, high, and max— which developers can configure upon request. This "effort" control allows for balancing intelligence, latency, and costFor trivial queries, it doesn't make sense to spend maximum resources, while for critical tasks or complex projects, it may be worthwhile to give the model room to think more.
These options are particularly useful in environments where thousands of API calls are executed daily, such as European SaaS, ERP integrations, or internal platforms of large companies. Being able to decide how much effort the model puts into each request makes it easier to manage AI costs without sacrificing quality when it's truly necessary.
In parallel, the platform adds features of automatic context compaction These tools help maintain very long work sessions, preventing abrupt interruptions when the token limit is reached. This type of fine-grained management is often key when integrating advanced models into existing systems.
Agent teams and working on large codebases
In the area of software development, Anthropic has announced for Claude Code the arrival of the so-called agent teams currently under investigation. The idea is that, instead of a single agent advancing serially, multiple agents could be launched. several sub-agents who divide the work in parallel and then they coordinate results, like a small engineering team would.
This approach is particularly interesting when working with large repositorieswhere reviewing all the code linearly is inefficient. Sub-agents can handle different modules, search for cross-references, analyze how data travels through the system, or review automated tests, and then unify their findings into a coherent result.
Combined with the 1 million token windowThis system allows, at least in theory, uploading an entire project and asking the model to locate weaknesses, inconsistencies, or risk areas without having to break down the files one by one. For development teams in Spain or the EU managing complex products, this can save days of manual review.
Opus 4.6 also strengthens its capabilities to code review and debuggingwith an emphasis on the model's ability to detect and correct its own errors during the process. This behavior, if it becomes established in everyday use, reduces the time developers have to spend reviewing code suggested by AI.
Benchmarks: where Claude Opus stands out 4.6

Anthropic supports the launch of Claude Opus 4.6 in a series of standardized tests where the model achieves outstanding results compared to its direct competitors. One of the most cited is Terminal-Bench 2.0, focused on agentic programming and autonomous work in console environments, where the model is at the top of the ranking.
En Humanity's Last ExamIn a multidisciplinary reasoning test that combines complex questions from various areas, Opus 4.6 also ranks ahead of other cutting-edge models, according to data provided by the company itself. This type of assessment aims to measure AI's ability to handle problems that more closely resemble real-world situations than isolated exercises.
Another relevant metric is GDPval-AA, which evaluates performance in high-value professional tasks in fields such as finance and lawIn this case, it is indicated that Opus 4.6 outperforms the next best model on the market —identified as GPT-5.2— by about 144 Elo points, and improves upon its own predecessor by an even greater margin.
The model also shows solid results in BrowseCompA test focused on locating information that is difficult to find online. Although benchmarks should always be taken with a grain of salt, the dataset points to a particularly strong model in Professional work, complex analysis, and precision searches, key areas for European banks, consultancies, law firms or technology companies.
Excel, PowerPoint and the leap to office work
Opus 4.6 doesn't just focus on the code: Anthropic has strengthened integrations geared towards office work, an area where AI is beginning to become an everyday tool for non-technical professionals. In the case of ExcelThe new version allows you to handle longer and more complex tasks without losing consistency, including advanced features such as Conditional formatting, data validation, or multistep changes on large sheets.
The model is capable of planning its actions before executing them, which helps prevent cascading errors when working with workbooks full of interconnected formulas. Furthermore, it can handle unsorted data, infer reasonable structures and prepare the ground for further analysis or visualizations.
The other big new feature is the integration of Claude in PowerPointCurrently in preliminary phase and primarily aimed at users of Max, Team, and Enterprise plans, this version may Generate complete presentations from a simple promptEdit existing slides and create charts while maintaining templates, fonts, and styles set by the organization.
In practice, this means that tasks such as preparing a monthly report, revising a sales deck, or updating a board presentation can be supported by AI without leaving the tool itself. For many European companies accustomed to the Microsoft ecosystem, this type of integration is more tangible than any benchmark.
Security, cybersecurity, and usage controls
Every leap in AI model capability is accompanied by doubts about security and misuseAnthropic has dedicated a significant portion of the Opus 4.6 announcement to this point. The company states that the new model maintains, and even improves upon, the alignment levels of previous versions, with low rates of problematic behaviors such as deception, excessive flattery, or cooperation for dangerous purposes.
A reduction in calls is also noted “excessive negativity”These are cases where a model refuses to answer harmless questions out of excessive caution. For professional environments operating under European regulatory frameworks, finding the balance between security and usability is a delicate matter, especially in regulated sectors such as finance or healthcare.
In the matter of cybersecurityAnthropic acknowledges that a more powerful model can both help defend and attack systems, and claims to have developed new internal probes to detect potentially harmful responses. At the same time, it is encouraging the use of the model to Identify and correct vulnerabilities in open source software, reinforcing the defensive angle.
The company also mentions the application of broader assessment batteries, including tests focused on user well-being, resistance to dangerous requests, and the detection of covert behavior. While much of this data comes from proprietary documentation, it is relevant for any company considering integrating Opus 4.6 into sensitive processes.
Pricing, availability, and deployment options

On the economic front, Claude Opus 4.6 It remains in the same price range as its predecessor, an important detail for companies that had already run the numbers on Opus 4.5. The standard cost is set at $5 per million entry tokens y $25 per million exit tokens, which puts the model in line with other high-end systems on the market.
For uses exceeding the 200.000 context tokensAnthropic applies a "premium" fee of approximately $10 per million tokens entered and $37,5 per million tokens exitedDesigned for projects that make the most of the enlarged window. Furthermore, the model supports up to 128.000 exit tokensThis is relevant when the expected result is a lengthy report, a large refactoring, or a set of documents generated all at once.
Opus 4.6 is available in claude.aivia Proprietary API and through the main cloud platforms, including those from Amazon and Google, which facilitates their adoption in enterprise environments already built on these infrastructures and in projects of local AIFor charges that require residency or data sovereignty requirements, the option of processing exclusively in US data centers. with an approximate surcharge of 10% on the price per token.
Regarding access plans, some advanced features—such as PowerPoint integration or certain intensive use modes—are restricted to higher subscription levels (Max, Team, and Enterprise), while users with free accounts remain on lighter models like Haiku or Sonnet. This aligns with the product's increasingly enterprise-focused approach.
What could this mean for companies and startups in Europe?
For tech startups, digital SMEs and large European corporations, Claude Opus 4.6 It arrives at a time of intense pressure to automate high-value tasks without triggering templates. The combination of Agentic programming, a 1 million token window, and tools like Excel and PowerPoint It positions the model as a serious option for rethinking entire workflows.
Small teams can rely on the model to assume projects that previously required more specialized profilesFrom building and maintaining complex codebases to preparing comprehensive reports for clients or regulators. For larger organizations, the advantage lies in integrating the model into internal systems and letting it take on some of the repetitive workload that currently falls to skilled professionals.
In terms of competition, the constant comparison with GPT-5.2 and Gemini 3 Pro This reflects a market in which several high-level offerings compete to become central to daily work. Each company will have to evaluate not only benchmarks, but also service conditions, data location, integration with their tools, and regulatory requirements, especially within the framework of the European AI Act.
By combining massive context, fine-tuned reasoning control, security improvements, and a clear focus on programming and office applications, Claude Opus 4.6 positions itself as a candidate for those seeking an AI that can handle long projects and demanding tasks without behaving like a simple demonstration chatbot.
With all of the above in mind, Anthropic's new model is shaping up to be another piece in an increasingly crowded AI offering, but with a clear angle: Supporting real professional work in coding, analysis, and documentation, maintaining relatively stable prices and adding the control functions that both developers and business managers expect in an increasingly demanding European regulatory environment.
I am a technology enthusiast who has turned his "geek" interests into a profession. I have spent more than 10 years of my life using cutting-edge technology and tinkering with all kinds of programs out of pure curiosity. Now I have specialized in computer technology and video games. This is because for more than 5 years I have been writing for various websites on technology and video games, creating articles that seek to give you the information you need in a language that is understandable to everyone.
If you have any questions, my knowledge ranges from everything related to the Windows operating system as well as Android for mobile phones. And my commitment is to you, I am always willing to spend a few minutes and help you resolve any questions you may have in this internet world.
