Independent perspectives. A connected world.About the publication ↗
G↗GLOBALTECHRANKS.TECHNOLOGY IN PERSPECTIVE
AI

Claude Sonnet 4.6 Expands Coding and Computer-Use Capabilities

Anthropic released Claude Sonnet 4.6 on February 17, 2026, describing improvements in coding, computer use, planning and long-context work. The company also introduced a one-million-token context window in beta and made the model the default for Free and Pro users in Claude and Claude Cowork.

Claude Sonnet 4.6 Expands Coding and Computer-Use Capabilities

Anthropic released Claude Sonnet 4.6 on February 17, 2026, describing improvements in coding, computer use, planning and long-context work. The company also introduced a one-million-token context window in beta and made the model the default for Free and Pro users in Claude and Claude Cowork.

In its announcement, Anthropic said pricing remained at Sonnet 4.5 levels, starting at $3 per million input tokens and $15 per million output tokens. The capability claims come from the vendor's evaluations and early feedback, rather than independent testing by GlobalRanking.

A larger context is a larger input budget

A long context window can help an application include more material in a request. That does not establish that the model will identify every important fact or apply every constraint correctly. Capacity and reliable use of that capacity are different evaluation questions.

For a development team, the useful test is a task that actually requires the additional information. A repository change might depend on conventions scattered across several files. An office task might need the contents of several documents. The evaluation should check whether the result reflects the relevant material, rather than merely whether the request was accepted.

It should also account for the cost of preparing and reviewing the input. Sending more information can reduce repeated work in one situation and add unnecessary processing in another. A larger allowance is an option to use deliberately.

Computer use brings operational decisions

The ability to interact with existing software is attractive when an application lacks a convenient integration. It also means the model must interpret screens, distinguish similar controls and decide what to do when the expected page changes.

GlobalRanking's assessment is that a business should begin with tasks where the result is easy to inspect. Updating a test record or assembling a draft report can reveal behavior without making every mistake consequential. The task should include a defined point where the user takes over.

The evaluation also needs an unexpected case. An expired login, incomplete form or conflicting instruction on a page can change how much human attention the workflow requires. A smooth demonstration should not be treated as evidence that those cases are solved.

Compare the completed task, not the model name

A model upgrade can improve useful work without improving every task equally. Coding teams should compare reviewed changes, tests and the effort needed to correct a result. Office teams should assess accurate outputs and the time required to validate them.

Keeping the same representative tasks across versions makes that comparison more informative. It also lets a team see whether an improvement in one area introduced a regression elsewhere.

Sonnet 4.6 gives Anthropic customers another model to evaluate at a familiar price tier. The practical decision remains whether it completes their work accurately enough, with a manageable amount of supervision and a cost that reflects the entire workflow.

Image: Anthropic

← Back to the latest