Skip to content

Open Weights and AI Leadership: How Boards Decide Whether Their Company Owns Its Intelligence

Open Weights and AI Leadership: How Boards Decide Whether Their Company Owns Its Intelligence

It feels almost impossible to keep up.

Every week there is a new “breakthrough” model, a new agent feature, a new voice mode, another trillion parameters, another million‑token context. It is exciting and exhausting at the same time. Just as boards and executives start to understand one release, another arrives. Into that blur comes the latest entry: Claude Opus 5.

Claude Opus 5 is Anthropic’s clearest sign that near‑frontier intelligence is meant to sit inside everyday work, not just in the research lab. It is now the default model on Claude Max and the strongest option on Claude Pro, and it performs close to Anthropic’s own Fable‑class systems while keeping the same price as Opus 4.8. In practical terms, a level of reasoning and coding that felt “frontier” only months ago is now priced and positioned as the new workhorse. Recent coverage has stressed the same point. Anthropic is presenting Opus 5 as its best‑performing and most cost‑effective model, explicitly designed to be used every day at half the price of Fable 5, in a market where enterprises are fretting about AI costs and demanding clearer links between spend and return.

On Anthropic’s benchmarks, Opus 5 comes within half a percentage point of Fable on CursorBench at roughly half the cost per task, beats Fable on OSWorld at just over one‑third of the cost, and scores three times higher than the next‑best model on ARC‑AGI 3. It more than doubles Opus 4.8’s performance on Frontier‑Bench while staying at about $5 per million input tokens and $25 per million output tokens. That combination of capability and cost makes it realistic to use Opus 5 all day for software, analytics and complex problem‑solving.

It is also built to handle a wide range of work. Anthropic positions Opus 5 as a primary choice for coding, data analysis, design, biology and knowledge tasks. It is tuned for building and debugging software, exploring datasets, iterating on product or interface designs, and supporting life‑science workflows that mix literature review, modeling and analysis. That breadth means Opus 5 can genuinely act as a default engine across multiple functions.

The way it behaves is changing expectations too. In one test, the model was asked to recreate a machine part as a 3D model without being allowed to view the drawing. Rather than guessing or refusing, it wrote its own computer‑vision pipeline, pulled the geometry out of raw pixels, and then rebuilt the part. That is the kind of step you would expect from a strong engineer: build the missing tool, then solve the problem. Models are starting to take that kind of initiative inside your workflows.

Security and reliability are evolving alongside capability. Anthropic’s internal evaluations suggest Opus 5 is their least “prompt‑injectable” model to date. Prompt injection is the class of attack where instructions hidden in input data try to trick the model into ignoring its rules, leaking secrets or taking unsafe actions. Opus 5 combines stronger alignment, dedicated prompt‑injection probes and a guarded “Auto Mode” in Claude Code that checks each tool call before it runs. In constrained coding and agent environments, layering those defenses has driven the success rate of common prompt‑injection patterns close to zero. That does not remove risk, but it changes the baseline for using Opus 5 as a trusted default agent.

All of this is happening inside a dense release cadence. Anthropic shipped Opus 4.8 on May 28 and then released Mythos 5, Fable 5 and Sonnet 5 in June, with Haiku still waiting for its own 5‑series upgrade. Google introduced Gemini 3.6 Flash together with 3.5 Flash‑Lite and 3.5 Flash Cyber on July 21, turning 3.6 Flash into a long‑context workhorse with better coding, built‑in computer use and lower API prices. OpenAI made GPT‑5.6 (Sol, Terra, Luna) its new flagship family on July 9. Moonshot AI’s Kimi K3 and Thinking Machines’ Inkling pushed the open‑weight frontier with models that carry trillions of parameters, million‑token contexts and multimodal inputs, both arriving in mid‑July. Even SpaceXAI’s Grok 4.5 entered that same frontier band for coding and agent work at aggressive prices.

Technical terms like “2.8 trillion parameters” or “one‑million‑token context” are easy to ignore. A more tangible view is to treat parameters as dials on a control panel and context as the amount of information the model can hold in mind. Models with trillions of dials can fit finer patterns in your code and documents. Models with book‑length memory can treat a portfolio, a case file or a reporting pack as one continuous story instead of a stack of files.

At the same time, the main labs are competing to become the way people interact with all of this. OpenAI’s GPT‑5.6 rollout now includes GPT‑Live voice models and a desktop ChatGPT Voice mode, so people can speak multi‑step instructions, drive agents and let the app see their screen. Anthropic has updated Claude’s voice mode so Opus, Sonnet or Haiku can work through Gmail, Calendar, Slack, Notion and Canva, and added “Record a skill,” which learns workflows from short screen recordings. Voice and screen‑watching agents are becoming standard ways to drive these models.

The net effect is that frontier‑level intelligence is now fast, accessible and deeply embedded in tools teams already use every day. The question for leaders is no longer whether to adopt these systems, but how to keep control as they become part of everyday operations.


Open Weights as a Leadership Strategy

Beneath this race, NVIDIA is supplying much of the compute that makes training and serving possible. With the Nemotron 3 family, the company is showing that its ambitions go far beyond just selling chips. It wants to lead at the model and agent layers as well, and it is doing so with open weights.

Nemotron 3 comes in Nano, Super and Ultra variants. Nemotron 3 Ultra is the largest and most capable, an open‑weight Mixture‑of‑Experts model with roughly 550 billion total parameters and about 55 billion active at once, meaning it routes each request to a small group of specialist sub‑models instead of using the entire network every time. It is designed for multimodal input, long context windows and agentic tasks that run over many steps. Independent benchmarking places Nemotron 3 Ultra near the top of US open‑weight models on general intelligence measures, ahead of other American open systems, though still a notch behind the strongest open models from China.

NVIDIA’s technical material describes Nemotron 3 as built for agents. The models support extended contexts, combine transformer and newer architectures, and are tuned to coordinate multiple tools and agents rather than just answer single prompts. They are intended to be the backbone for systems where teams of AI “colleagues” work together on complex workflows.

Jensen Huang’s public stance reinforces this direction. He has argued that every serious company will need its own AI stack and that the future of AI will be both proprietary and open. The Nemotron 3 family and NVIDIA’s open‑model announcements show that in practice. NVIDIA wants to supply GPUs, open‑weight models, training data and libraries so that developers can build their own agents and specialized systems on top of NVIDIA hardware.

The Open Weights and American AI Leadership letter fits this picture. It argues that national leadership in AI will depend on building a strong open ecosystem that reaches every sector, with open weights as a core element, rather than relying only on a few closed frontier labs. At the same time, it is signed by organizations that stand to benefit from such an ecosystem, including NVIDIA with Nemotron 3. The letter is both a policy vision and a product signal.

For boards and executives, the national framing has a direct parallel inside the firm. The question becomes whether the company will practice “open weights and AI leadership” internally, by owning more of its intelligence stack, or whether it will rely mainly on rented capability from closed providers.


Renting Intelligence vs Owning It: What Boards Must Decide

Opus 5, Gemini 3.6, GPT‑5.6, Kimi K3, Inkling, Grok 4.5 and Nemotron 3 Ultra all promise more powerful AI. They also bring a set of governance decisions into focus. Boards now have to decide how much intelligence the company will own, and how much it will simply rent.

Relying chiefly on closed frontier models is a rental strategy. Systems send tokens to APIs from Anthropic, OpenAI, Google, SpaceXAI and others. The company receives high‑quality outputs, benefits from vendor support and uses polished features such as voice, agents and “recorded skills.” The company pays per million tokens and per user. For some customer‑facing and low‑sensitivity use cases, leveraging hardened, well‑supported cloud services can be appropriate. But in many regulated and high‑risk domains, relying mainly on closed, external models creates long‑term risks around data control, auditability and dependency.

If that approach quietly becomes standard everywhere, dependence grows. When most code, documents, analytics and workflows are mediated by one or two closed vendors, their decisions begin to shape the business. Their model updates change how systems behave. Their safeguards determine what agents are allowed to do. Their pricing and usage limits influence which processes can be automated and which cannot. Boards may find, a few years in, that they are watching core capabilities drift outside the organization without ever choosing that outcome.

Owning intelligence requires a different posture. It begins with treating data, ontology and the enterprise knowledge graph as core assets, and building an “intelligence stack” around them. Closed and open models then plug into that stack. It involves running capable open‑weight or self‑hosted models such as Nemotron 3 Ultra, Kimi K3 or Inkling on the company’s own infrastructure, tuned to its domain. It means designing systems so workloads can be routed across Opus, Gemini, GPT‑5.6, Grok, Nemotron and other models without re‑architecting everything above them, and so that the most sensitive and regulated workflows stay on models the company controls.

Open weights are what make this practical. When an organization runs open‑weight models, it can keep sensitive data inside its perimeter, inspect and log behavior more thoroughly, and build self‑improving capabilities that remain part of its asset base rather than the vendor’s. The trade‑offs around safety and misuse are real, but they can be managed with careful scoping, access control, red‑teaming and continuous evaluation. The gain is sovereignty over a growing share of the intelligence layer, which matters most where regulation, security and long‑term accountability are strongest.

For boards, this is not just a technology question. It touches strategy, risk and investor expectations.

Over the last two years, many companies have told investors that AI will drive growth, productivity and new revenue. Those narratives have appeared in earnings calls, investor days and slide decks. The time when those promises will be tested is now. If boards treat AI as a purely operational choice left to management, they risk drifting into AI‑washing: ambitious commitments without corresponding systems, metrics or returns.

To avoid that, boards can take several steps.

First, they can define where AI oversight lives. One primary committee, often audit or risk, should own the risk framework for AI: vendor dependence, open‑weight exposure, regulated workloads, emerging threats such as prompt injection and how model behavior is monitored and reported. A technology or innovation committee can oversee architecture, posture toward open versus closed models and progress against the AI roadmap. The full board should review major AI investments, strategic partnerships and public commitments.

Second, boards can ask management to map workloads into “rent” and “own” categories. That means identifying which tasks will rely on closed frontier systems, which will run on open‑weight or self‑hosted models, and which regulated, high‑sensitivity workflows must be kept on models the company controls. This mapping should cover both current deployments and planned expansions.

Third, boards can encourage practical moves toward owning more intelligence where it matters most:

Anthropic and OpenAI are already seeing the shift to a more cost‑conscious customer base, where enterprises are reluctant to keep testing expensive frontier models without a clear path to measurable returns. Boards should assume that this scrutiny will extend to their own AI narratives. If a company’s AI story leans heavily on high‑priced closed models, without a concrete plan for value creation, use of open weights or real control over its intelligence stack, AI spending can quickly turn into a visible cost line with no matching performance line - and investors will eventually call that out.

The pace of change is impossible to keep up with. These past few weeks we've seen one release after the other, shifting what “normal” looks like for intelligence in the enterprise. By the time directors finish reading an article about the newest model, the next model is already here, with more parameters, more features, and more ways for you to spend money.

This is the new normal. This is the tempo in which governance decisions now have to be made. The choice between renting and owning intelligence, between relying on closed APIs and investing in open weights, between telling a bold AI story and delivering measurable results, these are all things the boardroom has to grapple with in this new era of continuous governance of an increasingly agentic enterprise.

While boards can't control the pace of AI releases, they can control how they react to it. Companies that treat AI as a passing feature will be carried along by whatever their vendors ship. Companies that frame AI as an intelligence layer they intend to own - shaped by open weights, sovereign data and clear accountability - will have a chance to turn the exponential intelligence curve we all are on to their advantage. The work now is to build the strategic, technical, financial and ethical that let your company stay in control of its own intelligence as that curve keeps steepening.

Steven Wolfe Pereira

Steven Wolfe Pereira

Steven Wolfe Pereira is Founder & CEO of Alpha, an AI governance intelligence company serving boards and executives. Former C-suite executive at Datalogix / Oracle, Neustar, and Quantcast; board member, startup advisor and Forbes contributor.

All articles

More from Steven Wolfe Pereira

See all