# Part 2: The Agent Asked Once. The Workflow Ran Twice.

*In Part 1 I told you to expose your most boring Logic App as an MCP tool. It took five minutes. Then an agent got impatient, called it again, and I spent the rest of the week finding out what "production-ready" actually costs.*

*(Client work, so names, numbers and a few details are changed. The shape of it isn't.)*

So I did the thing I told you to do.

I took our most boring Logic App, the one that looks up an order in Azure SQL and hands back a JSON blob, and I ran it through the MCP wizard. Pointed Copilot at it. Asked it about an order. It found the tool, called it, and gave me the right answer on the first try.

I was very pleased with myself. I showed the team.

Then one of the seniors leaned back and asked, "Okay. Who else can call that?"

And I sat there, because the honest answer was: anyone with the key. The key that was sitting in my `mcp.json`. The key I had, *ahem*, pasted into a Teams chat so someone else could try it.

That question is this whole post. In [Part 1](https://dreaded-developer.hashnode.dev/your-logic-apps-were-already-agentic-you-just-didn-t-know-it) I said the portal wizard gets you to *working* and that production was a different job. This is that job: APIM in front, Entra instead of keys, a registry so people know which tools are fair game, and the governance bits I promised last time (one of which I got wrong, and I'll own that further down).

And the one that actually hurt. The day an agent asked for something once, and the workflow did it twice.

Fair warning, this one has more XML than I'd like. Integration life.

* * *

## What I'm actually building

Here's where we're going before we get into the weeds.

![Production stack: AI agent calls APIM with an OAuth2 token, APIM calls Logic Apps Standard with managed identity, slow work goes to Service Bus, API Center lists approved MCP servers, Entra handles identity](https://cdn.hashnode.com/uploads/gql/6a3beb0f1ccbe2c3cf93fc89/5faa0b20-b56c-4f88-aebe-3326e41d34da.png align="center")

In plain words:

- **The agent** (Foundry Agent Service in my case, but Copilot or Claude Desktop work the same way) never talks to the Logic App directly. It only knows the APIM URL.
- **APIM** is the front door. It checks who you are, how often you've been knocking, and whether what you're sending looks sane. Then it passes the call along.
- **The Logic App** is still the same workflow from Part 1. Nothing about the workflow changes. What changes is that it now only accepts calls from APIM.
- **Service Bus** sits behind the tools that take longer than an agent is willing to wait. More on that later, because that's where agents and integrations really don't get along.
- **API Center** is the list of MCP servers we've agreed people are allowed to use. Think of it as the menu, not the bouncer.
- **Entra** decides who gets in. Every agent gets an identity, and every identity gets only the roles it needs.

If you've built anything integration-heavy on Azure, none of these boxes are new. That's kind of the point of this whole arc. The only new thing is the caller, and the caller happens to be a language model that will happily try things you didn't expect.

## Putting APIM in front (and making sure nobody goes around it)

APIM can sit in front of an MCP server that already exists, which is exactly what a Logic App with the MCP server turned on is. In the portal it's **APIs → MCP Servers → Create MCP server → Expose an existing MCP server**. You give it the Logic App's MCP endpoint (copy it from the **MCP servers** blade on the Logic App rather than guessing the path), pick Streamable HTTP, and APIM hands you a new URL. That new URL is the only one your agents should ever see.

Three things I wish someone had told me before I started clicking:

- **Check your tier.** MCP servers work on Developer, Basic, Standard, Premium and the v2 tiers. Not on Consumption, and not inside workspaces. Our shared dev instance was Consumption. Of course it was.
- **Don't touch the response body in policy.** No `context.Response.Body` in MCP server policies, and set the frontend "number of payload bytes to log" to 0 at the All APIs scope. Both force APIM to buffer the response, and buffering breaks streaming. The symptom is an agent that connects fine and then just... hangs. Fun afternoon.
- **The backend has to speak a recent MCP spec** (2025-06-18 or later). The Logic Apps MCP server does. Your hand-rolled one from last year might not.

Now the part people skip. Putting APIM in front doesn't stop anyone calling the Logic App directly. If the `azurewebsites.net` URL still answers the internet, your gateway is decoration. A very well-configured piece of decoration.

My first instinct was an access restriction on the Logic App using the `ApiManagement` service tag. Don't. That tag covers APIM's *management* traffic, the control plane, not the requests your gateway sends out. And the v2 tiers don't have a fixed outbound IP to allowlist anyway.

What actually worked for us was two layers:

1. **Network.** A private endpoint on the Logic App, public access turned off, and APIM integrated into the VNet so it can reach that private endpoint. Now there's simply no public route to the workflow.
2. **Identity.** APIM calls the Logic App with its own managed identity, and the Logic App only accepts tokens from that identity. Which brings us to the key.

## Killing the key

The demo used a key. Keys are fine for a demo, and terrible for anything else, because a key says nothing about *who* is calling. It just says someone, somewhere, copied it.

The Logic Apps MCP server supports OAuth through App Service authentication (Easy Auth), alongside the API key. So the setup ends up with two separate token checks, one at each hop:

- **Agent → APIM.** Each agent gets its own identity in Entra. We registered one app for the gateway, exposed an app role per kind of tool (`Tools.Read` for lookups, `Tools.CreditMemo` for the stuff that moves money), and assigned roles to agent identities one by one. The support agent got both. The reporting agent got `Tools.Read` and nothing else.
- **APIM → Logic App.** APIM swaps the agent's token for its own, using its managed identity. The Logic App's Easy Auth only allows APIM's identity as a client, and we turned the API key off entirely.

One thing that tripped me up: the Logic Apps docs tell you to set Easy Auth's restrict-access option to **Allow unauthenticated access**. That looks wrong the first time you see it. It's because the MCP endpoint enforces the token check itself rather than letting the platform block the request. Which means you should actually test it with no token and a wrong token, and see a 401, before you trust it.

Here's the policy on the MCP server in APIM:

```xml
<policies>
    <inbound>
        <base />
        <!-- Who are you? Only tokens issued for our gateway app, carrying a tool role. -->
        <validate-azure-ad-token tenant-id="{{tenant-id}}" output-token-variable-name="agentToken">
            <audiences>
                <audience>{{mcp-gateway-app-id-uri}}</audience>
            </audiences>
            <required-claims>
                <claim name="roles" match="any">
                    <value>Tools.Read</value>
                    <value>Tools.CreditMemo</value>
                </claim>
            </required-claims>
        </validate-azure-ad-token>
        <!-- How often? Per agent, keyed on the calling app's client ID. -->
        <rate-limit-by-key calls="30" renewal-period="60"
            counter-key="@(((Jwt)context.Variables["agentToken"]).Claims.GetValueOrDefault("azp", "unknown"))" />
        <!-- Then call the Logic App as APIM, not as the agent. -->
        <authentication-managed-identity resource="{{logicapp-mcp-app-id-uri}}" />
    </inbound>
    <backend>
        <base />
    </backend>
    <outbound>
        <base />
    </outbound>
    <on-error>
        <base />
    </on-error>
</policies>
```

The `{{...}}` bits are named values, so nothing tenant-specific lives in the policy itself. The `azp` claim is the calling app's client ID on v2 tokens. If your tokens are v1, it's `appid`.

You'll notice the role check here is coarse. It lets in anyone who has *either* role. Making `Tools.CreditMemo` mandatory for the credit memo tool specifically means a per-tool check, and I'll admit we did it the boring way: the credit memo workflow lives in its own MCP server in the Logic App, behind its own MCP API in APIM, with its own policy that requires that one role. More boxes, less cleverness. I can live with that.

And because "it compiles in the portal" isn't a test, here's the small console app I keep around to poke the gateway as a real caller. It gets a token the same way an agent would, lists the tools, and then calls the credit memo tool *twice with the same key*. You'll see why that second call matters in a minute.

```csharp
using System.Text.Json;
using Azure.Core;
using Azure.Identity;
using ModelContextProtocol.Client;
using ModelContextProtocol.Protocol;

// Smoke test: call a side-effecting tool twice with the same key, through APIM,
// with a real Entra token. If the two answers differ, the tool isn't safe for agents.
var gatewayUrl = new Uri(Environment.GetEnvironmentVariable("MCP_GATEWAY_URL")
    ?? throw new InvalidOperationException("Set MCP_GATEWAY_URL"));
var scope = Environment.GetEnvironmentVariable("MCP_GATEWAY_SCOPE")
    ?? throw new InvalidOperationException("Set MCP_GATEWAY_SCOPE, e.g. api://<app-id>/.default");

var credential = new DefaultAzureCredential();
AccessToken token = await credential.GetTokenAsync(new TokenRequestContext([scope]));

var transport = new HttpClientTransport(new HttpClientTransportOptions
{
    Endpoint = gatewayUrl,
    Name = "credit-memo-smoke-test",
    AdditionalHeaders = new Dictionary<string, string>
    {
        ["Authorization"] = $"Bearer {token.Token}"
    }
});

await using McpClient client = await McpClient.CreateAsync(transport);

IList<McpClientTool> tools = await client.ListToolsAsync();
Console.WriteLine($"Gateway exposes {tools.Count} tool(s): {string.Join(", ", tools.Select(t => t.Name))}");

var arguments = new Dictionary<string, object?>
{
    ["caseId"] = "case-5512",
    ["invoiceNumber"] = "INV-20931",
    ["amount"] = 1250.00m,
    ["idempotencyKey"] = "case-5512"
};

string first = await StartCreditMemoAsync(client, arguments);
string second = await StartCreditMemoAsync(client, arguments);

Console.WriteLine($"first:  {first}");
Console.WriteLine($"second: {second}");
Console.WriteLine(first == second
    ? "OK: same correlationId, the retry was absorbed."
    : "FAIL: two different correlationIds. An agent retry would do the work twice.");

static async Task<string> StartCreditMemoAsync(McpClient client, Dictionary<string, object?> arguments)
{
    CallToolResult result = await client.CallToolAsync("start_credit_memo", arguments);

    if (result.IsError == true)
    {
        throw new InvalidOperationException("start_credit_memo returned an error.");
    }

    string json = result.Content.OfType<TextContentBlock>().Single().Text;
    using JsonDocument doc = JsonDocument.Parse(json);
    return doc.RootElement.GetProperty("correlationId").GetString()
        ?? throw new InvalidOperationException("No correlationId in the tool result.");
}
```

That's a .NET 10 console app with two packages, `ModelContextProtocol` (2.2.0) and `Azure.Identity`. Run it under the agent's identity, a managed identity or a service principal, not your own account. Your account probably has more roles than the agent does, and the whole point is to see what the *agent* sees.

## The agent asked once. The workflow ran twice.

Once the gateway and the Entra bits were in, we let a small pilot loose: a customer service agent for the support team, with a handful of Logic App tools behind it. Look up an invoice. Check a shipment. And the one that matters for this story, `issue_credit_memo`.

The workflow behind it was nothing fancy. It's the kind I've built a dozen times since the BizTalk days. Validate the request against Azure SQL, post the credit memo to the ERP through the on-prem data gateway, render the PDF, email the customer. On a normal day, twenty to forty seconds end to end. A human support rep had already approved the credit in the chat before the agent ever called the tool, so we felt pretty good about the guardrails.

Then month-end close happened.

If you've done ERP integration, you already know. During close, finance is hammering the ERP with reports and postings, and everything that touches it slows to a crawl. Our "twenty to forty seconds" workflow was sitting at the ERP step for three minutes and change.

The agent's tool call didn't wait that long. It gave up at a hundred seconds and got back an error. That number isn't a setting, by the way. For non-streaming MCP tool calls, Foundry Agent Service just stops waiting at 100 seconds, and I learned that from run history, not from the docs. Meanwhile the Logic App was perfectly happy to hold the request for 225 seconds, the Standard default, before giving up on the caller itself. From where the agent sat, the tool had failed. So it did what any reasonable caller does with a failed call. It tried again. It even told the rep, very politely, "Sorry, that didn't go through the first time. I've retried and the credit memo has been issued."

It had. Twice.

The first run never stopped. A stateful Logic App run doesn't care that the caller hung up; it just keeps going. So the first run finished its ERP posting, and a little under two minutes later the second run finished its own. Two credit memos, same invoice, same amount. Two emails to the customer, too, which is how a customer finds out you have a bug before you do.

![Sequence diagram: the agent calls issue_credit_memo, the call times out while the workflow is still waiting on the ERP, the agent retries, and both runs post a credit memo for the same invoice](https://cdn.hashnode.com/uploads/gql/6a3beb0f1ccbe2c3cf93fc89/496cfc22-4c7a-4333-86af-0ffebb4cdecf.png align="center")

We actually found it two days later, when someone from AR pinged us during reconciliation: "Why does this customer have two credit memos for INV-20931?" I opened run history and there they were. Two runs, identical inputs, both green. *Both green.* That was the part that got me. Every dashboard we had said everything was fine.

It wasn't just the one, either. Three customers got double credits that afternoon, all during the slowest hour of close. Finance reversed them, nobody lost their job, and I lost a weekend.

Here's what bugs me about it. Nobody did anything wrong, exactly. The agent behaved like a well-built client: the call failed, so it retried. The workflow behaved like a well-built workflow: it got a valid request, so it did the work. The bug lived in the gap between them. The agent can't tell a *failed* call from a *slow* one, and the tool had no idea it had already been asked.

And the most embarrassing part? In Part 1, in the Honest Take, I wrote "Design your Logic Apps to be idempotent. Use correlation IDs. Handle duplicate calls gracefully." I wrote that. Then I exposed a workflow that moves money without doing any of it, because it had never needed to. Its old callers were a scheduled batch and a web form with a disabled submit button. Neither of them retried on their own. *Ayun.* An agent does.

## Making slow tools safe to call twice

There are two fixes, and you want both. The first stops most retries from happening. The second makes the ones that still happen harmless.

![Sequence diagram: start_credit_memo enqueues work on Service Bus with the caller's idempotency key as MessageId and returns a correlation ID; a repeated call is dropped by duplicate detection; the workflow checks the ERP for the key before posting; the agent polls get_credit_memo_status](https://cdn.hashnode.com/uploads/gql/6a3beb0f1ccbe2c3cf93fc89/f9a1905a-f37b-420b-9bd9-9b11002e537e.png align="center")

**Fix one: stop making the agent wait.** `issue_credit_memo` is gone. In its place are two tools and a worker:

- `start_credit_memo` is a tiny workflow. Request trigger, validate the input, drop a message on a Service Bus queue, return `202` with a correlation ID. It's done in about a second, no matter how sad the ERP is feeling.
- A separate workflow, triggered by the queue, does the real work: ERP posting, PDF, email. It writes its progress to a status table in Azure SQL as it goes.
- `get_credit_memo_status` reads that table. The agent polls it, and the tool description literally says "call this every 30 seconds until status is `done` or `failed`; do not call start_credit_memo again". Models are surprisingly good at following that when you spell it out.

This is the oldest pattern in integration, by the way. It's the async request-reply pattern every BizTalk person has drawn on a whiteboard a hundred times. The only new thing is that the client is a model. (If you're wondering whether Durable Functions would make this easier: it has the pattern built in, and I went through when I'd pick it over Logic Apps in [Logic Apps vs. Durable Functions](https://dreaded-developer.hashnode.dev/logic-apps-vs-durable-functions-what-i-wish-i-knew-before-choosing).)

**Fix two: make it idempotent anyway.** Fix one shrinks the window, but it doesn't close it. The agent can still call `start_credit_memo` twice. A network blip, a model that gets confused, a rep who asks again. So:

- **The key comes from the business, not the model.** The idempotency key is derived from the support case ID. Never let the model invent the key. A model that makes up a fresh GUID on every retry has just made your idempotency check useless, very confidently.
- **The correlation ID comes from the key.** Same key in, same correlation ID out. That's what the smoke test checks: call it twice, get the same answer.
- **Service Bus drops the obvious duplicates.** The key becomes the message's `MessageId`, and the queue has duplicate detection on. Two catches here. You can only turn it on when you *create* the queue, so yes, we recreated ours. And it needs Standard or Premium; Basic doesn't have it. The window defaults to 10 minutes and goes from 20 seconds up to 7 days.
- **The worker checks before it posts.** This is the one that actually saves you. The worker stamps the key into the credit memo's external reference field in the ERP, and before posting it looks for an existing memo with that reference. If one exists, it marks the request done and stops. The queue's dedupe window is a seatbelt. This check is the brakes.

I want to be clear that none of this is clever. It's the stuff I'd have demanded in a design review for any other caller. We skipped it because the old callers never retried, and "it's just exposing an existing workflow" made it feel like we weren't really building anything new. We were. We were adding a new caller with completely different manners.

## Telling people which tools are allowed

Once there's more than one MCP server, people start asking which ones they're allowed to use. Usually they ask after they've already wired one into their editor.

We use API Center as the private MCP registry. Register the APIM-fronted MCP server there (the APIM URL, never the Logic App one), and you get a catalogue people can browse, plus a registry endpoint that VS Code and GitHub Copilot can point at. On the GitHub side, an org or enterprise admin can set the MCP policy to **Registry only**, so Copilot only offers servers from your list.

Two honest caveats:

- GitHub still labels "Registry only" as public preview, and its own docs say it isn't the recommended way to *restrict* access. It's a strong nudge, not a wall.
- To make the registry readable by those clients, you'll likely turn on anonymous read access in API Center. Which means the menu itself isn't a secret. It shouldn't be, but think about what you put in the tool descriptions.

So, API Center is the menu, not the bouncer. The bouncer is still Entra plus APIM. A registry tells well-behaved clients where to go. It does nothing about a badly behaved one that already knows an address, which is why the private endpoint and the token checks from earlier matter more than this section does.

## About that "token governance" I promised

In Part 1 I promised "the content safety and token governance policies you need" on the MCP gateway. Half of that promise was wrong, and I'd rather own it than quietly skip it.

APIM's `llm-token-limit` policy counts *language model tokens*: prompt and completion, read from the model's usage data, on APIs that speak an LLM schema like OpenAI's chat completions. An MCP tool call to a Logic App doesn't spend model tokens. There's nothing for the policy to count. I'd mixed up two different gateways in my head.

Here's how I'd split it now:

- **On the model endpoint** (the Foundry or Azure OpenAI deployment your agent thinks with, ideally behind APIM as well): `llm-token-limit`. Tokens per minute per agent, which returns 429 when exceeded, and a quota per period, which returns 403. This is where the money actually goes, and it's the same territory as the cost guardrails in [Part 4 of the agentic AI arc](https://dreaded-developer.hashnode.dev/part-4-agentic-ai-in-production-blue-green-deployments-circuit-breakers-and-cost-guardrails).
- **On the MCP tool endpoint:** `rate-limit-by-key` per agent, like the policy above, and `quota-by-key` if you want a daily ceiling. An agent stuck in a loop calling your lookup tool 400 times a minute is the failure you're guarding against here, not token spend.
- **On both:** `llm-content-safety`, which can now check MCP tool requests and responses too, not just prompts. It sends the payload to Azure AI Content Safety and blocks it with a 403 if it's over your thresholds. Useful if your tools pass free text from customers straight into the agent's context, which ours do.

## The Honest Take

**An agent is the least patient, most retry-happy client you will ever expose a workflow to.** It doesn't know your ERP has a bad week every month. It doesn't know a failed call and a slow call are different things. It will retry, and it will tell the user it worked. If a tool has a side effect, it has to be safe to call twice. Full stop. If you take one thing from this post, take that.

**"Just expose the existing workflow" is the trap.** The workflow was correct for its old callers. Exposing it to an agent is a new integration with a new client, and it deserves the same design review you'd give any other new client. We skipped the review because the wizard made it feel like a settings change.

**Check what's actually GA before you build on it.** In Part 1 I called the Logic Apps MCP server generally available. As I write this, the Learn page for it still carries the preview banner. Either the announcement got ahead of the docs or I got ahead of both. Either way, treat it as preview: watch for breaking changes, and don't make it the only path to anything critical. The APIM side is on firmer ground.

**Start with zero tools exposed, and make each one earn its spot.** Every MCP tool is a new public API with a very creative caller. Our rule now: read-only tools need a named owner. Anything with a side effect needs a named owner, an idempotency key from the business, and a passing run of that smoke test. Most workflows don't need to be agent-callable at all, and that's fine.

**The registry is the least important box in the diagram.** It's useful, and I'd still set it up, but it's the one people will point at when security asks "how do we control which tools agents use?" The real answer is identity and network. Don't let a nice catalogue page stand in for either.

## So what do you do this week?

Open the list of workflows you've exposed as MCP tools, or the ones you're planning to. For each one with a side effect (it posts, sends, creates, pays) ask one question: *what happens if this gets called twice, ninety seconds apart, with the same input?*

If the answer is anything other than "nothing, the second call returns the first call's result", that workflow isn't ready for an agent. Fix that before you fix anything else in this post.

* * *

## What's Coming in Part 3

Remember the part of the story that actually kept me up? Not the duplicate. The fact that both runs were green. APIM logged two successful calls, the Logic App logged two successful runs, and Service Bus would have happily logged two messages too. Every box did its job, and none of them could tell me that a single conversation with a single agent had caused all of it.

That's Part 3: being able to answer "what did the agent actually do?" after the fact.

- Following one agent conversation end to end, from the model's tool call through APIM and the Logic App to Service Bus and the ERP, with one correlation ID instead of four disconnected logs
- What APIM, Logic Apps and Application Insights each give you out of the box, and where the trail goes cold
- The alerts I wish we'd had during month-end close: same tool, same key, twice in five minutes
- Making run history useful to the support team and finance, not just to whoever built the workflow

*Ingat sa retries.* Watch out for the retries.

## References

- [Create MCP servers from workflows in Azure Logic Apps (Standard)](https://learn.microsoft.com/en-us/azure/logic-apps/create-model-context-protocol-server-standard)
- [Limits and configuration for Azure Logic Apps (HTTP request timeouts)](https://learn.microsoft.com/en-us/azure/logic-apps/logic-apps-limits-and-config)
- [Model Context Protocol tool in Foundry Agent Service (MCP tool call timeout)](https://learn.microsoft.com/en-us/azure/foundry/agents/how-to/tools/model-context-protocol)
- [Overview of MCP servers in Azure API Management](https://learn.microsoft.com/en-us/azure/api-management/mcp-server-overview)
- [Expose an existing MCP server in Azure API Management](https://learn.microsoft.com/en-us/azure/api-management/expose-existing-mcp-server)
- [Secure access to MCP servers in Azure API Management](https://learn.microsoft.com/en-us/azure/api-management/secure-mcp-servers)
- [validate-azure-ad-token policy](https://learn.microsoft.com/en-us/azure/api-management/validate-azure-ad-token-policy)
- [authentication-managed-identity policy](https://learn.microsoft.com/en-us/azure/api-management/authentication-managed-identity-policy)
- [rate-limit-by-key policy](https://learn.microsoft.com/en-us/azure/api-management/rate-limit-by-key-policy)
- [llm-token-limit policy](https://learn.microsoft.com/en-us/azure/api-management/llm-token-limit-policy)
- [llm-content-safety policy](https://learn.microsoft.com/en-us/azure/api-management/llm-content-safety-policy)
- [Azure service tags overview](https://learn.microsoft.com/en-us/azure/virtual-network/service-tags-overview)
- [Register and discover MCP servers in Azure API Center](https://learn.microsoft.com/en-us/azure/api-center/register-discover-mcp-server)
- [Configure an MCP registry for your organization (GitHub Docs)](https://docs.github.com/en/copilot/how-tos/administer-copilot/manage-mcp-usage/configure-mcp-registry)
- [Service Bus duplicate detection](https://learn.microsoft.com/en-us/azure/service-bus-messaging/duplicate-detection)
- [Asynchronous Request-Reply pattern (Azure Architecture Center)](https://learn.microsoft.com/en-us/azure/architecture/patterns/async-request-reply)
- [MCP C# SDK: transports](https://csharp.sdk.modelcontextprotocol.io/v2/concepts/transports/transports.html)

