To orchestrate multi-agent workflows efficiently, you can combine Agent-to-Agent (A2A) routing and a Model Router within Azure AI Foundry. This allows a primary Orchestrator Agent to delegate complex data operations to a specialized Data Extraction Agent (which calls your Azure Function MCP Server), while dynamically shifting traffic between gpt-4o (for reasoning) and gpt-4o-mini (for lower-cost data processing).
Step 1: Deploy and Define the Models in the Project
Before configuring the Model Router, you must ensure both target models are deployed in your Azure AI Foundry hub.
- Go to the Azure AI Foundry Portal (azure.com).
- Under Shared Resources, navigate to Models + Endpoints.
- Deploy two models if you haven’t already:
gpt-4o(Name the deployment:gpt-4o-heavy)gpt-4o-mini(Name the deployment:gpt-4o-light)
Step 2: Create the Sub-Agent (Data Extraction Agent)
The downstream Sub-Agent will explicitly handle interacting with your custom Azure Function MCP server to fetch file paths.
We initialize this agent using the cheaper model (gpt-4o-mini) because structured tool execution does not require deep reasoning.
python
import os
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import Agent, AgentToolbox
# 1. Initialize Project Client
project_connection_string = os.environ["AZURE_AI_PROJECT_CONNECTION_STRING"]
client = AIProjectClient.from_connection_string(
credential=DefaultAzureCredential(),
conn_str=project_connection_string
)
# 2. Reference your existing Function App MCP Tool
# (Assuming it's already cataloged in your Foundry workspace)
mcp_tool_id = "your-azure-function-mcp-tool-id"
# 3. Create the Specialized Downstream Agent
data_extraction_agent = client.agents.create_agent(
model="gpt-4o-mini", # Using the low-cost deployment for data fetching tasks
name="Data-Extraction-SubAgent",
instructions=(
"You are a specialized file path extraction agent. Your only job is to use the "
"provided Azure Function MCP tool to fetch structured JSON file paths from Blob and SharePoint. "
"Always return the clean JSON output back to the Orchestrator."
),
tools=[{"type": "mcp", "id": mcp_tool_id}]
)
print(f"Sub-Agent Created Successfully. ID: {data_extraction_agent.id}")
Step 3: Bundle the Sub-Agent into a Toolbox (A2A Setup)
To allow a primary Orchestrator Agent to call this sub-agent, the sub-agent must be wrapped into an Agent-to-Agent (A2A) tool and bundled inside a Foundry Toolbox.
python
# 4. Wrap the Sub-Agent as an A2A Tool definition
a2a_tool = {
"type": "agent_to_agent",
"settings": {
"agent_id": data_extraction_agent.id,
"description": (
"Use this tool when the user requests file paths, file listings, "
"or data ingestion tasks from Azure Blob Storage or Microsoft SharePoint."
)
}
}
# 5. Create or Update a versioned Foundry Toolbox with this A2A Tool
toolbox = client.agents.create_toolbox(
name="EnterpriseRoutingToolbox",
description="Toolbox containing data extraction sub-agents and routing policies.",
tools=[a2a_tool]
)
print(f"Toolbox published with A2A capabilities. URI: {toolbox.mcp_endpoint_uri}")
Step 4: Implement the Model Router and Orchestrator Agent
A Model Router evaluates incoming requests and forwards them to the appropriate model based on complexity. For multi-agent orchestration, we will set up the primary agent to use a dynamic routing profile: it boots on gpt-4o for high-level semantic planning but routes sub-tasks out to the toolbox efficiently.
python
from azure.ai.projects.models import ModelRouterConfig, ModelRouterRule
# 6. Define Model Routing Rules (JSON-schema style routing)
# If a prompt contains phrases explicitly asking for simple status checks or raw lists,
# the platform can automatically downgrade the primary LLM to save token costs.
routing_config = ModelRouterConfig(
default_deployment="gpt-4o-heavy",
rules=[
ModelRouterRule(
condition="prompt.contains('status check') or prompt.contains('just list')",
target_deployment="gpt-4o-light"
)
]
)
# 7. Create the Orchestrator Agent utilizing the Model Router and A2A Toolbox
orchestrator_agent = client.agents.create_agent(
model_router=routing_config, # Dynamic cost optimization
name="Primary-Orchestrator",
instructions=(
"You are the main coordinator. Analyze user prompts. If a user asks to scan, "
"retrieve, or process files from storage systems, delegate the task completely to the "
"Data-Extraction-SubAgent tool. Synthesize its output back to the user."
),
toolbox_id=toolbox.id # Injecting the Toolbox containing the A2A sub-agent connection
)
print(f"Orchestrator configured with Model Router and A2A Toolbox. Ready for ingestion.")
Step 5: Execute the Multi-Agent Flow
When a user asks a complex question, the primary orchestrator processes the reasoning via gpt-4o, identifies that it needs storage paths, routes execution natively via A2A to the sub-agent, which subsequently runs the python codebase on your Azure Function App MCP Server.
python
# Create a user thread
thread = client.agents.create_thread()
# Submit a complex multi-layered request
client.agents.create_message(
thread_id=thread.id,
role="user",
content="Find all CSV sales reports modified yesterday in SharePoint and give me a summary report."
)
# Run the Orchestrator
run = client.agents.create_run(thread_id=thread.id, agent_id=orchestrator_agent.id)
# Process the stream
while run.status in ["queued", "in_progress"]:
run = client.agents.get_run(thread_id=thread.id, run_id=run.id)
# Print execution transcript
messages = client.agents.list_messages(thread_id=thread.id)
for msg in reversed(messages.data):
print(f"[{msg.role.upper()}]: {msg.content[0].text.value}")
Use code with caution.
Step 6: Verify Telemetry in Application Insights
Because both agents are bound to the Azure AI Project client workspace, your Application Insights dashboard will capture a cascading hierarchy trace:
- Top-Level Span:
Primary-Orchestratorexecution viagpt-4o-heavy. - Child Span (A2A Route): Event showing delegation to
Data-Extraction-SubAgent. - HTTP Webhook Dependency: Post request flowing straight into your
MyFoundryMcpAppAzure Function endpoint.
To monitor this routing overhead and cost tracking in App Insights Logs, run this query:
kusto
customEvents
| where name in ("AgentExecution", "ToolCall", "ModelRouting")
| extend ModelUsed = tostring(customDimensions.model_deployment)
| project timestamp, name, ModelUsed, duration
NEXT: enforce custom Input/Output Guardrails at the Orchestrator level to block sensitive system paths from being sent down to the sub-agent, or we can configure a Shared Thread Memory pattern so both agents can access a history of files pulled previously.