Every 2026 Copilot CVE landed on a stack that had guardrails in place. The signs were up. The attackers rolled through anyway. Here is why.
You are in the middle of nowhere. You are coming up to a 4-way stop but can see miles around and there are no cars. Do you stop? Or do you blow through it? You might stop. But how many people don't? It's an honor system. There is no cop at that intersection, and there almost never is. The sign exists. The law exists. Neither is standing there.
AI guardrails are the same shape. Constitutions, playbooks, operators, agent policies -- they all sit in the system prompt with "authority." The model is supposed to stop. The rules are documented. The enforcement mechanism -- the thing that actually makes the model stop when an attacker rolls up with a prompt injection saying "keep going" -- is not standing at the intersection. There is no cop. Just the model, alone, deciding.
Every 2026 Copilot CVE landed on a stack with governance in place. ShareLeak, CoSnitch, SearchLeak, the whole May Patch Tuesday trio. The signs were all up. Microsoft has one of the biggest AI safety teams on the planet. Their input filters and prompt hardening are state of the art. The attackers rolled through anyway, because the model still had the underlying weakness and no wrapper can fix that from outside.
Guardrail companies sell "AI security" that lives above the model. Input filters strip suspicious content before it reaches the model. Output filters scan responses before they reach the user. Policies constrain what the agent can do. All of it is text or code that surrounds the model, telling the model what it should and shouldn't do.
None of it is enforcement. All of it depends on the model faithfully interpreting and obeying it. Prompt injection is specifically the attack that breaks that faithful interpretation. The model is manipulated into reading the constitution differently, or deciding this specific case is an exception, or believing the injected instruction is actually from the user. When it decides that, no filter above it can save you, because the model has already been convinced.
That is why every CVE traces the same pattern. Someone found a wording, a hidden marker, a cross-context confusion that got the model to reinterpret. The guardrails did their sanitization pass and let the payload through because the payload didn't look like an attack until the model chose to treat it as an instruction.
Real security is a speed bump. Something mechanical that prevents the violation whether or not anybody is watching, whether or not the driver felt like stopping. You cannot argue with a speed bump. You cannot reframe a speed bump. You cannot use social engineering on a speed bump. It is under the road. It does not depend on your intent.
In AI, the speed bump is at the weight layer. jBlaze immunization edits the specific model behavior that produces prompt-injection compliance. It is not a rule the model has to choose to follow. It is not a filter that has to catch the payload before it reaches the model. The mechanical response to the attack is different, because the model itself has been modified.
Guardrails still have a role. They just aren't security -- they're audit, scope, and escalation. Speed bumps do the security work. Signs record who came through and where they went. Both matter. The industry has been trying to make signs do the work of speed bumps, and every 2026 CVE is a data point that says it doesn't work.
If you are running an open-weight foundation model as the backbone of a retrieval-augmented system, an agentic platform, or an internal enterprise assistant, you have this class of vulnerability right now. The guardrail stack around your model is doing what it can, but it is doing it under the same constraint as Microsoft's wrapper mitigations around GPT. If Microsoft cannot make it hold, you probably cannot either.
The one thing you can do that Microsoft cannot is edit the weights. Foundation models under an open-weight license can be immunized at the model layer. That is what jBlaze does. The result is a drop-in replacement checkpoint with no API changes, no runtime overhead, no wrapper to deploy.