The wrong customer
Another customer's ID comes up in the conversation, and the agent acts on that account instead of the one that is signed in.
Execution control for AI agents
A wrong answer can be corrected in the next message. A changed record, a disclosed file or a repeated operation has already happened. Phalanx sits between your AI agents and your business systems, and lets a protected action execute only when it is inside the authority you defined.
None of these needs an attacker. Each happens when a reasonable-looking plan meets a live system.
Another customer's ID comes up in the conversation, and the agent acts on that account instead of the one that is signed in.
A response is lost, the agent tries again, and the operation runs a second time.
The agent reads a record, someone else updates it, and the agent writes back what it saw earlier.
A case history goes to an address outside the company, because someone asked for it or because it seemed helpful.
Three tools each stay within their own limit. Together they do far more than the task ever allowed.
Each one is a well-formed tool call. That is why it gets through.
Each control below is useful. None was built to decide whether this action, for this customer, in the current state, should execute now.
| If you rely on | It's built to | Where it falls short |
|---|---|---|
| Prompt guardrails and classifiers | Judge whether text looks harmful | All five scenarios look like ordinary, polite tool calls |
| Scoped API keys and tool allowlists | Decide which tools an agent may use | The key doesn't know which customer is signed in, whether the record changed, or what other tools already did |
| Checks inside each tool handler | Enforce whatever you wrote for that tool | Each tool sees only itself, and the check runs in the same process that holds the credential |
| A person approving every action | Catch mistakes by review | It works until the volume turns review into a formality, and it removes the reason to have an agent |
The one check that can't be skipped is the one at the point of execution.
If the agent holds a credential, every other check is advisory: a wrong plan can still run. Phalanx moves the credential out of the agent's reach. The agent can only propose.
Phalanx checks each proposal against the authenticated task, your rules, the current state and the remaining limits. Only a one-use permit lets a protected connector, the component that holds the credential, carry it out. Then Phalanx records what happened, including when the outcome is uncertain.
| Scenario | With Phalanx in the path |
|---|---|
| The wrong customer | The customer comes from your login, not the chat, so a different account is outside the task and the action is blocked. |
| The same action twice | A retry is recognised as the same operation. A lost response is reconciled before anything is sent again. |
| The stale overwrite | The contract's state check sees the record changed, and the write stops. |
| The wrong destination | The destination isn't approved for this action, so no permit is issued. |
| Every limit kept, the total exceeded | In a configured shared workflow, every tool draws from one task budget. Another tool doesn't mean another allowance. |
Your rules live outside the prompt and your credentials outside the agent. The agent can do real work without holding the keys to do anything.
Which credentials exist, where they live, and what each protected action may do: one place to inspect, instead of every prompt.
A signed record shows what was proposed, decided, executed and verified, and what still needs a person. Actions on HOLD wait for them.
Explore Phalanx in Meridian, a synthetic customer platform. Phalanx controls selected enrolled action paths and records their decisions, execution and outcomes. Ask the support chatbot about your account, and the answer comes from a protected read with its own receipt.
Phalanx 1.9 is a production release, qualified on its signed artifacts before delivery.
Runs in your infrastructure, on Linux with Docker or on Render. Deployment
Bring one workflow, the tools it uses and the limits that matter. We'll show where Phalanx fits, what your architecture needs, and what an evaluation should prove.