Back to Blog

Endpoint Hygiene Was Never an IT Problem. It's a Budget Problem in an IT Costume.

By: Casey Cannady : nomad, cybersecurity veteran & Chapter 7 survivor

August 12, 2026
8 min read
Casey Michael Cannady
CybersecurityEndpoint ManagementPersonal Story

TL;DR

Yesterday Microsoft shipped fixes for 398 vulnerabilities in a single Patch Tuesday. That is not the story. The story is the trend line: roughly 200 in June, more than 570 in July, 398 in August, with Microsoft openly crediting AI-assisted vulnerability discovery for the flood. Adobe has moved to twice-monthly bulletins. Cisco, Google, Mozilla, and Oracle are all shipping more, faster. This is not a spike. It is the new floor. And the uncomfortable part for anyone still running patch management out of a spreadsheet and a prayer is that AI is turning out to be excellent at finding holes and mediocre at closing them, which means the human-plus-tooling side of this equation just became the whole ballgame. I have spent nearly 30 years on that side, most of it in enterprise endpoint management, a good chunk of it as an IBM Certified BigFix guy. Full disclosure lives in its own section below, because I want you reading the argument, not wondering who paid for it. Short version: I have been platform agnostic since December 2024 and I am not selling you a product here. I am telling you that the absence of centralized endpoint management is no longer a technical debt. It is an unfunded liability, and it is almost never IT's fault.

They will hand you the invoice for a decision you were never allowed to make. Then they will ask why you did not patch faster.


The Weekend I Got To Go Home

WannaCry hit while I was at Kroger.

You remember the week. Ransomware chewing through hospitals and rail networks and anything with an unpatched SMB stack, executives refreshing news sites, security teams sleeping under desks. Every organization I knew was in a war room trying to answer one question they should have already known the answer to: are we exposed?

We were not, and more importantly, we could prove it in near real time.

That was not luck and it was not heroics. It was the Kroger Endpoint Management team, KREM, having already done the boring, unglamorous, chronically underfunded work of getting a management agent onto essentially every supported compute endpoint in the enterprise. Desktops. Laptops. Servers. The whole sprawling footprint of a nationwide grocery chain with well over a hundred thousand endpoints scattered across stores, distribution centers, plants, and offices.

Because that inventory existed, the answer to “are we exposed” took minutes instead of days. Because the answer took minutes, a lot of people who expected to lose their weekend got to go home.

Nobody wrote a case study about it. That is the tragedy of endpoint hygiene. When it works, the story is that nothing happened, and nothing happened is the hardest thing in the world to put in a budget request.


What Actually Happened Tuesday

Brian Krebs' write-up of yesterday's Patch Tuesday is the immediate reason I am writing this one. Go read it.

The headline number is 398 vulnerabilities, 42 of them rated critical. The sole known zero-day is a privilege escalation flaw in afd.sys, the driver sitting underneath Windows socket connections on effectively every endpoint you own.

But the number is not the point. The sequence is:

  • June 2026: roughly 200 fixes. A record at the time.
  • July 2026: more than 570 fixes. A new record.
  • August 2026: 398 fixes.

Microsoft has attributed the deluge to vulnerability discovery aided by artificial intelligence. Adobe now publishes security bulletins twice a month. Cisco, Google, Mozilla, and Oracle have all increased volume and cadence.

Sit with what that means operationally. Your patch backlog is no longer an event you respond to. It is weather you live in. Any process built on the assumption that a monthly patch cycle is a discrete, human-reviewable batch of a few dozen items has already failed. It just has not been told yet.


The Zero Day Is Step Two

Here is the detail from yesterday's release that I cannot stop chewing on, because it connects directly to everything I wrote in my last series.

Automox's Landon Miles described that afd.sys zero-day as “step two in a chain.” Step one is a phish. An attacker talks a human into giving up a low-privilege foothold, and only then does the driver flaw get used to take the whole box.

I spent three posts earlier this year hammering one point from three different angles: nothing gets hacked, something gets trusted that should not be. Russian intelligence did not break Signal's encryption, they impersonated support and asked nicely. A fake AI agent skill did not defeat any scanner, it borrowed a star count. The lock keeps holding. Somebody keeps talking the keyholder into opening it.

You cannot patch the human. I have run security awareness programs for decades and I will tell you plainly that you can move the needle, you cannot close the gap.

So if step one is permanently available to the attacker, your entire defensive posture rests on step two. How fast can you see the foothold, how small is the blast radius around it, and how quickly can you close the escalation path? Every one of those questions is an endpoint management question. Not an EDR question, not a firewall question, not a training question. Can you see the machine, do you know what is on it, and can you change it.


AI Finds. Humans Fix.

There is a comfortable story going around that AI is about to eat this problem, and organizations that are behind on endpoint hygiene are quietly betting on it. I want to be careful here, because I am not an AI skeptic. I build with these tools every single day and they make me faster at nearly everything I do.

But the evidence in Krebs' piece deserves your attention. Researchers at 1Password tested what happens when large language models generate patches for newly disclosed, complex vulnerabilities. More than half the time the resulting patch either failed to fix the flaw, introduced a new weakness, or both.

Ed Skoudis at the SANS Technology Institute reports genuinely good results using AI for patch generation, with one condition: humans in the loop, testing and iterating. His summary is that it still needs a skilled human at the keyboard.

That is exactly right, and it is exactly where I land. AI belongs in endpoint management. Triage, correlation, drafting remediation content, summarizing a 398-item release into what actually applies to your fleet, all of it. What AI does not do is own your change control policy, understand why your point-of-sale image is special, or know that the last time you touched that driver on that hardware model you took down four hundred stores.

Automate the analysis. Keep a qualified human accountable for the process. Do both, or you get neither.


Nobody Said Rush. Everybody Needs To See.

Tyler Reguly at Fortra made a point in the Krebs piece that runs against every vendor email currently sitting in your inbox: of nearly 400 bugs, exactly one is known to be actively exploited. His advice was blunt, something to the effect that there's no need to rush these updates, regardless of what vendors tell you.

He is right, and I want to draw the line under why that is not a contradiction of everything else in this post.

The ask is not speed. The ask is visibility.

You cannot triage 398 items down to the twelve that matter in your environment if you do not know what is in your environment. You cannot deploy in staged rings if you have no rings. You cannot pause a rollout at 5% when telemetry goes sideways if you have no telemetry. And you cannot tell your board “we are not exposed” if the honest answer is “we think most of them are patched.”

Organizations without central management do not patch slowly. They patch blindly, and then they rush, which is how you get the second kind of outage.


The War Room Where the Deployment Tool Is Guilty Until Proven Innocent

Let me give you the other side of the ledger, because I am not going to pretend this discipline is free.

Years back at Kroger, the desktop engineering group pushed an antivirus upgrade across the fleet. Routine. Tested. Approved. It went out through BigFix, because BigFix was the deployment mechanism for everything, which meant that when problems started, I got pulled into the war room. When you own the delivery pipe, you are guilty until proven innocent. Every endpoint management professional reading this just nodded.

What we eventually untangled was this: the AV upgrade had an undocumented side effect that forced the endpoint to fully re-cache Active Directory and DNS. One machine doing that is nothing. Thousands of machines doing it in a coordinated wave is a self-inflicted denial of service against your own directory and name resolution infrastructure. And the waves lined up precisely with the other team's staggered deployment schedule.

The deployment tool was innocent. The deployment tool was also the only reason we could correlate the waves to the rollout and stop it.

The lesson I took out of that room and have repeated for years since: small-scale testing hides production-scale side effects. A hundred-machine pilot cannot surface a problem whose entire mechanism is aggregate load. Which is why “do not rush” and “deploy in rings with real-time reporting” are not two pieces of advice. They are the same piece of advice.


It Was Never an IT Failure

Here is the part I actually came to say.

In nearly 30 years of walking into organizations at every scale, from a nationwide grocery chain to a U.S. federal agency running north of 300,000 endpoints that I am not permitted to name, I have almost never found an IT team that did not know what good looked like. I have found a great many that were never funded to build it.

The pattern is so consistent it is practically a law:

  1. Endpoint management is proposed. It is unglamorous, it produces no new revenue, and its entire value proposition is the absence of future catastrophe.
  2. It gets cut, deferred, or funded at a fraction that covers licensing but not the staff to run it. Everyone involved understands this is a gamble. Nobody writes that down.
  3. The incident arrives.
  4. The incident is described, in the postmortem and in the press, as an IT failure.

That last step is the con. The people who declined to fund the levee get to stand on the roof and ask why the levee did not hold. Regular readers know I am not alleging a conspiracy. I am describing architecture. You do not need bad actors to produce this outcome. You just need a budget process that discounts invisible risk and an accountability process that lands on the people furthest from the checkbook.

If you are a CISO, you already know this. Your job is making the gamble legible before it pays out. Get the risk acceptance in writing, signed, at the level that actually made the call. Not to cover yourself, though it will. To force the decision to be a decision instead of a drift.


Full Disclosure

I owe you this before the next section, because the next section has opinions about money.

I spent a large chunk of my career as an IBM Certified BigFix expert, including six and a half years at HCL Software after the product moved there. I was frequently brought into projects specifically because of infrastructure and scale experience. I know that platform down to the bone, and I have a real affection for what it does well.

I resigned in 2024. Part of the reason was that Karen and I wanted to start 3D Nomadic and get on with building 3D-printed nomadic habitats. The other part is that after six and a half years, I was denied a cost of living adjustment, having demonstrated that I was earning less in 2024 than in 2018 once you adjusted for inflation. That is the sound the old deal makes when it breaks, and I have written about it elsewhere.

Since December 2024 I have been deliberately platform and tool agnostic. I do not resell anything. I do not have a partner agreement with anyone. I currently have consulting work in flight and an existing client relationship continuing, none of it tied to a vendor. If I name products below, it is as existence proof for an argument, not a recommendation, and I want that stated plainly rather than buried.


The Per-Endpoint Meter

Now, the money.

Nearly every commercial endpoint management and security platform is licensed per endpoint, per year, frequently with bundled minimums and multi-year commitments. There is a legitimate business logic to that. There is also a well-documented pattern where the price per seat climbs at every renewal, features you already depend on get repackaged into a higher tier, and the cost of leaving is deliberately engineered to exceed the cost of absorbing the increase.

I have written about this shape before and given it a name borrowed from Benn Jordan: Leveragism. You do not own the capability. You rent access to it. And the meter never stops.

I sat in those meetings on the vendor side. Nobody in them uses the word trap. They say stickiness.

If you run an MSP, this is not abstract, it is your margin. You resell a per-seat license across your client base. The vendor raises the rate at renewal. Now you have exactly two options: eat the increase and watch your margin compress on every single seat you manage, or pass it through to a client who will absolutely use that invoice as a reason to get three competing quotes. The vendor takes the raise. You take the churn risk. That is not a partnership, that is a toll booth with a logo on it.

Here is what has genuinely changed in the last few years, and it is the most useful thing in this post: the open source options are no longer a compromise.

Wazuh, osquery and Fleet, Ansible, OpenSCAP, and the broader ecosystem around them are running in serious production environments doing serious work. On the commercial side, BigFix, Tanium, Microsoft Intune, Ivanti, and Automox all have real strengths and real installed bases. I am not ranking any of these. I am telling you the category is no longer “expensive real tools versus free toys,” which is a sentence I could not have written with a straight face fifteen years ago.

The trade is honest and you should go in clear-eyed. Commercial buys you support contracts, vendor accountability, and someone to call at 3 a.m. Open source buys you no per-endpoint meter, full data ownership, and no renewal negotiation, in exchange for staffing the expertise yourself. That is a real trade with real costs on both sides. What it is not, anymore, is an excuse. “We could not afford endpoint management” stopped being true, and the organizations still saying it are describing a priority, not a price.


I Refuse to Leave You in the Dark

Same split as always. Structural fixes that actually move the needle, and things you can do Monday. Do not confuse the two.

The levee (what we should be demanding):

  • Fund hygiene before the incident, in writing. Endpoint management is not an IT line item. It is operational risk reduction and it should be defended at the same level as insurance.
  • Make fleet coverage a board-visible metric. Not “are we patched.” What percentage of known endpoints checked in within the last 24 hours. Everything else is downstream of that one number.
  • Procurement that does not punish open source. Policies requiring a vendor support contract by default disqualify capable tools on paperwork grounds, not merit.
  • Contract terms that assume you will leave. Exportable data in open formats, no punitive exit, price escalation caps written in. If leaving is impossible, you are not a customer, you are a tenant.
  • Name the decision-maker in every postmortem. If the gap traces back to an unfunded request, that belongs in the report as prominently as the CVE does.

The sandbags (what you can do Monday):

  • Build the inventory, even badly. A spreadsheet you actually maintain beats a platform you never deployed. You cannot manage what you cannot enumerate, and step one is always enumeration.
  • Deploy in rings, always. Pilot, then a representative slice, then the fleet. Production-representative, not just production-sized.
  • Assume aggregate behavior differs from individual behavior. My AV war room in one sentence. If a change touches directory, DNS, authentication, or anything shared, stagger it and watch the infrastructure, not just the endpoint.
  • Keep the human in the AI loop. Use it to triage 398 items down to your twelve. Do not let it own the change policy.
  • Get the risk acceptance signed. If the funding is not there, document what that buys and who bought it. Today, not after.

Be clear-eyed, though. Rings and spreadsheets and clever triage are sandbags. They help you survive the quarter. They do not build the levee. Only funded, staffed, organizational commitment does that.


Why I'm Telling You This

Because I have watched too many good technical people absorb the blame for a financial decision they were not in the room for.

I am not writing this as a vendor, and after this year I am definitely not writing it as someone with a comfortable relationship to institutional power. I filed Chapter 7 last October. I have spent the months since building a company that prints homes on wheels, which is about as far from an enterprise patch window as you can get and still be working with your hands.

But I keep coming back to the same shape, whether it is the middle class, a gas pump, a grill brush, or a fleet of 300,000 endpoints. Somebody makes a decision that trades tomorrow's catastrophe for this quarter's number, and then the catastrophe lands on somebody else's name.

The trend line in that Patch Tuesday post is not going to reverse. AI is going to keep finding vulnerabilities faster than humans can responsibly close them, and the gap between organizations that can see their fleet and organizations that cannot is about to become the single most consequential dividing line in operational security. That is not a prediction. It is arithmetic.

Despair says the volume is impossible and you were always going to drown. Fury says this was a funding choice made by people who can be named and who are, for the most part, still in the room, which means it can be chosen differently. Despair files the postmortem. Fury builds the levee.

Go read Krebs' piece. Then go find out how many of your endpoints checked in yesterday. And get loud with me.


Sources & Further Reading

  • Krebs on Security: “Microsoft Plugs Nearly 400 Security Holes” (August 11, 2026), the source for the 398 and 42-critical figures, the June and July patch counts, Microsoft's attribution to AI-aided discovery, and the Automox, 1Password, SANS, and Fortra commentary quoted here.
  • Automox: Landon Miles' August 2026 Patch Tuesday analysis, source of the afd.sys zero-day breakdown and the “step two in a chain” framing.
  • 1Password: research on why AI-generated vulnerability patches still require human review.
  • SANS Internet Storm Center: per-patch breakdown by severity and urgency for August 2026, the practical companion to any of this.
  • Prior posts in this vein: “They Didn't Break Signal. They Broke the Person Holding the Phone.” on why the human is the load-bearing part of every system, and “It Was Never a Dream. It Was a Subscription.” on Leveragism and the ownership-to-access conversion this post applies to security tooling.

Sourcing note: every hard figure here comes from the Krebs article linked above, which I cross-checked against the live post before publishing. The Kroger and government-scale anecdotes are my own experience, told without figures or identifiers where I am not free to provide them. Verify anything you plan to repeat.


Connect with Casey

If this resonated, or if there's a topic you want me to take on next, reach out. I read everything.

Casey writes about economic policy, nomadic life, cybersecurity, and navigating the world as a late-diagnosed AuDHD adult. New posts drop on my professional website.