The Software Is Starting to Do the Work

 

The Software Is Starting to Do the Work

Agentic AI, what it changes, and why we’re pursuing Anthropic partner status.

What Actually Changed

For the last few years the useful version of AI in business was a very good assistant. You asked it something and it answered: a draft, a summary, a block of code you still had to place, test, and own. The work stayed with you.

What shipped over the past year is different in kind. An agent takes a goal instead of a question. It reads the actual codebase, runs the actual query, opens the ticket, checks its own output against a real result, and comes back when it is finished or when it needs a decision from a person.

The distinction is not how smart it sounds. The distinction is whether it can act, and whether it stops in the right places.

That second half is the entire engineering problem. A system that can act on your data and your infrastructure is either carefully curated or it is a liability, and nothing in the model itself decides which one you get.

Where It Belongs, and Where It Doesn’t

The unpopular half first, because it is the half that decides whether an engagement works.

Plenty of the problems we get called for are not AI problems. They are an unmaintained system nobody wants to touch. A manual process that has never been written down, and so cannot be automated, because nobody can say precisely what it is. Four tools doing one job because each was bought in a different year by a different person. Selling a model into any of those would be easy and wrong. You would pay for AI and still have the original problem, now with a layer on top of it.

Genuinely an AI problem
â—ŹHigh-volume work with a clear definition of done
â—ŹWork that spans systems never designed to talk to each other
â—ŹReview and triage a person should approve rather than perform
●Long-tail tasks never worth a developer’s week
Not an AI problem
Ă—An unmaintained system nobody wants to touch
Ă—A manual process that has never been written down
Ă—Four tools doing one job bought in four different years
×Anything where “correct” cannot be defined up front
Figure 1 · The sorting test

Where agentic systems genuinely earn their keep is narrower, and more useful, than the marketing suggests. High-volume work with a clear definition of done. Work that spans systems that were never designed to talk to each other. Review and triage passes that a person should approve rather than perform. Long-tail tasks that were never worth a developer’s week but cost a department a day every month.

The test we use is simple. If you can describe what “correct” looks like, an agent can probably do it and you can probably check it. If you cannot, no model is going to work that out on your behalf.

What We Already Ship With It

We build with AI, and we are not going to be shy about it. By 2026 every serious shop is using these tools (or soon will be), and the ones claiming otherwise are usually just not saying so. It is part of why our client’s ten-phase claims system went from discovery to beta in about nine weeks.

The nine weeks matters less than what surrounds the generation.

01
Agent proposes
the change
→
02
Version control
and human review
03
Tested against
your real data
→
04
Documented as
the work happens
Repeat until it meets the spec — nothing unreviewed reaches production
You own the code, the data, and the infrastructure — during the engagement and after it.
Figure 2 · How agentic work actually ships
  • Version control and review on everything. No unreviewed generated code reaches your production environment.
  • Tested against your real data, not a demo fixture. Systems that look right in a walkthrough and fall over on the actual dataset are the single most common way software projects fail.
  • Documentation written as the work happens, not reconstructed at handoff.
  • Your repository, your infrastructure. The work product lives where you can reach it, during the engagement and after it.

Where Does the Data Go

Almost nobody asks this early enough.

Most software risk in 2026 is not a breach. It is a platform doing exactly what it was designed to do  calling home, enriching your records against a third-party service, syncing something you did not know was syncing  all of it documented somewhere in a settings page nobody opened. AI tooling makes that surface larger, not smaller.

So we audit the platforms we deploy for outbound behavior, including the ones we run our own business on. What talks to the outside, and what it sends. Which of those are on by default, because free and community-tier software frequently ships data-sharing enabled  that is the commercial model. Where the boundary is enforced, which belongs at the network layer and not only in an application checkbox a vendor update can quietly reset. And what we can turn off and still ship.

What that means for you is a written answer to where your data lives, what leaves your environment, how long it is retained, and what  if anything  it trains. If we cannot answer that about a component, we do not put your business on it.

Since 1993, We Have Seen/Done This Before

Founded
1993
Commercial
web
Mobile
Cloud
Distributed
ledger
Agentic AI
2026
Figure 3 · Waves we have built through

We have been building software for over thirty years, through the arrival of the commercial web, mobile, cloud, and distributed ledger. We were working with smart contracts and stable coins back when bitcoin was worth pennies.

We bring that up because thirty years of arrivals is what lets us tell a shift from a cycle.

Every one of those waves arrived with the same two claims attached: that it changed everything, and that you had to move immediately or be left behind. Both were about half true every time. The firms that came out ahead were the ones that learned the technology properly and then applied it where it actually fit.

Longevity also answers the most expensive risk in a custom software engagement. Buyers tend to worry about the build. What actually costs them is the firm disappearing two years in and leaving behind a system nobody can maintain. Thirty years is an answer to that, and very few firms in this market can give it.

Agentic AI is a real shift. We are treating it the way we treated the others.

Why We Are Seeking Claude Partner Network Status

SBLOCK is pursuing membership in the Anthropic Claude Partner Network. We are not members yet, and we are not going to describe ourselves as one until we are.

The program requires our team to work through Anthropic’s Academy learning path before status is granted. That is the work in front of us, and it is deliberately not quick. We would rather say so plainly than let a credential appear on this site before it is real.

We are doing it because “we use AI” is not a qualification. Every firm says it, and the phrase has stopped carrying weight. Training against the platform vendor’s own curriculum, and being held to their standard rather than to our own opinion of what good looks like, is a way of making the claim mean something. It also keeps us close to where the platform is going, instead of reading about it after it has already landed.

What Partner Status Would Mean

Stated conditionally, because it has not landed yet.

For Clients

  • A team trained to a published external standard, not only to our own judgment.
  • Earlier visibility into platform direction, so that what we recommend has a longer shelf life than the current release.
  • Vetted practice for the part that matters most — how your data is handled inside AI systems, which is the question that decides whether any of this is usable in a regulated or sensitive environment.
  • No change to what actually differentiates the engagement. Software, infrastructure, security, and media stay under one roof. Integration remains our problem, not yours.

For the Firms We Work Alongside

Referral and channel relationships get the same thing clients do, from the other side of the table: somewhere to send the AI-adjacent scope of an engagement without handing over the client relationship, and a partner who will tell you honestly when the work in front of you is not an AI problem at all.

Where This Leaves You

If someone is pitching you agentic AI right now, the questions worth asking are not about the model.

Ask what it is allowed to touch. Ask what happens when it is wrong. Ask where the data goes and who owns the result when the engagement ends. Ask who maintains it in three years.

We will give you those answers about our own work, including the parts where the answer is that AI is not what you need.

“SBLOCK successfully brought our company into the 21st century. I didn’t realize how much effort and resources we were wasting. Thank you!”

Michelle Chan, Procurement Manager

Talk to Us

We look over your whole environment — software first — and tell you what is worth doing, what is worth automating, and what should be left alone.

Request a Consultation
+1 (800) 407-6875

Software engineers since 1993. Technology partners to Orlando, Miami, and Palm Beach.

 

Claude Opus 4.7 Just Launched. Here’s What It Actually Changes for Business.

Anthropic released Claude Opus 4.7 today. If you’re running a business that depends on software, handles documents, or is evaluating AI tools — this one matters. Not because it’s the flashiest launch of the year, but because of what specifically improved and who it’s built for.

What Actually Changed

Claude Opus 4.7 is a direct upgrade to the Opus 4.6 model that powered the Knuth breakthrough we covered last month. The improvements are targeted, not cosmetic.

Software engineering got meaningfully better. Opus 4.7 scored +13% on a 93-task coding benchmark compared to its predecessor, and resolved 3x more production-level tasks on Rakuten-SWE-Bench. On CursorBench — which measures real developer workflows — it hit 70%, up from 58%. These aren’t toy benchmarks. They’re measuring whether the model can actually ship code.

It’s dramatically more efficient. In enterprise evaluations by Box, Opus 4.7 used 56% fewer model calls, 50% fewer tool calls, responded 24% faster, and consumed 30% fewer AI Units than the previous version. That translates directly to lower API costs for businesses running Claude at scale.

Document analysis improved substantially. On Databricks’ OfficeQA Pro benchmark, Opus 4.7 made 21% fewer errors when working with source documents — financial reports, contracts, technical specifications. For any business that processes paperwork, that’s a measurable reduction in mistakes.

Vision got a 3x resolution upgrade. The model now processes images at more than three times the resolution of Opus 4.6. Charts, dense documents, screen UIs, and slide decks are all handled with significantly higher accuracy. If you’ve ever pasted a screenshot into an AI chat and gotten a vague response, this is the fix.

Long-running tasks stay on track. Opus 4.7 delivered the most consistent long-context performance of any model tested, tying for the top overall score across six evaluation modules. For businesses running multi-step workflows — research, analysis, code generation, reporting — the model no longer drifts off course halfway through.

Why This Matters Beyond the Benchmarks

The numbers are strong, but the real story is about what kind of company Anthropic is becoming — and what that signals for businesses evaluating AI vendors.

Anthropic now has over 1,000 enterprise customers paying more than $1 million annually for Claude services. Their annual recurring revenue has hit $30 billion, and analysts project it could triple by year-end. Claude’s share of chatbot traffic nearly doubled between February and March 2026. This isn’t a research lab anymore. It’s a platform company with serious enterprise traction.

The UK government is using Claude to power GOV.UK, the country’s main public information portal. The British government is actively courting Anthropic for further expansion, including a potential dual stock market listing. When a G7 government selects your AI for citizen-facing services, that’s a credibility signal that matters.

Opus 4.7 is available everywhere businesses already deploy. It launched simultaneously on the Claude API, Amazon Bedrock, GitHub Copilot, Google Cloud, and Microsoft Azure. If you’re on any of those platforms, the upgrade is a configuration change — not a migration.

The Elephant in the Room: Mythos

CNBC reported today that Anthropic describes Opus 4.7 as their most powerful generally available model — but positions it as “less broadly capable” than Claude Mythos Preview, their unreleased frontier model. That distinction matters.

Mythos is the ceiling. Opus 4.7 is the floor that businesses can actually build on today. And for most real-world applications — writing code, analyzing documents, automating workflows, processing images — the floor just got raised significantly.

What This Means for Your Business

If you’re already using Claude, this is a free upgrade. Opus 4.7 is a drop-in replacement for Opus 4.6 across every deployment channel. You get better results at lower cost without changing a single line of integration code.

If you’re evaluating AI tools and haven’t committed yet, the landscape just shifted. The efficiency gains alone — 56% fewer API calls, 24% faster responses — change the unit economics of AI-powered automation. Projects that didn’t pencil out at Opus 4.6 pricing might work now.

If you’re a software team, the coding improvements are the headline. A model that resolves 3x more production tasks and scores 70% on real developer workflow benchmarks isn’t an assistant anymore. It’s a junior engineer that works around the clock.

And if you’re in an industry that runs on documents — legal, financial services, insurance, healthcare — the 21% error reduction in document analysis is the number to focus on. That’s not a marginal improvement. That’s the difference between an AI tool you have to babysit and one you can trust.

So What’s the Move?

The businesses that gain the most from a model release like this aren’t the ones that rush to adopt. They’re the ones that have already mapped out where AI fits into their operations and can slot the upgrade into an existing workflow.

If you haven’t done that mapping yet, that’s where SBLOCK comes in. We advise on AI tool selection, integration architecture, and automation strategy — for software teams, operations teams, and leadership trying to figure out which of these capabilities actually matter for their specific business.

The model got better. The question is whether your business is set up to take advantage of it.

Request a Consultation

SBLOCK has been building with Claude since the early access days. We know what it’s good at, where the limits are, and how to integrate it into production systems that have to work every day — not just pass a benchmark.

Request a Consultation