Setup the Agent-to-Agent (A2A) tool and model router

To orchestrate multi-agent workflows efficiently, you can combine Agent-to-Agent (A2A) routing and a Model Router within Azure AI Foundry. This allows a primary Orchestrator Agent to delegate complex data operations to a specialized Data Extraction Agent (which calls your Azure Function MCP Server), while dynamically shifting traffic between gpt-4o (for reasoning) and gpt-4o-mini (for lower-cost data processing).


Step 1: Deploy and Define the Models in the Project

Before configuring the Model Router, you must ensure both target models are deployed in your Azure AI Foundry hub.

  1. Go to the Azure AI Foundry Portal (azure.com).
  2. Under Shared Resources, navigate to Models + Endpoints.
  3. Deploy two models if you haven’t already:
    • gpt-4o (Name the deployment: gpt-4o-heavy)
    • gpt-4o-mini (Name the deployment: gpt-4o-light)

Step 2: Create the Sub-Agent (Data Extraction Agent)

The downstream Sub-Agent will explicitly handle interacting with your custom Azure Function MCP server to fetch file paths.

We initialize this agent using the cheaper model (gpt-4o-mini) because structured tool execution does not require deep reasoning.

python

import os
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import Agent, AgentToolbox

# 1. Initialize Project Client
project_connection_string = os.environ["AZURE_AI_PROJECT_CONNECTION_STRING"]
client = AIProjectClient.from_connection_string(
    credential=DefaultAzureCredential(),
    conn_str=project_connection_string
)

# 2. Reference your existing Function App MCP Tool
# (Assuming it's already cataloged in your Foundry workspace)
mcp_tool_id = "your-azure-function-mcp-tool-id" 

# 3. Create the Specialized Downstream Agent
data_extraction_agent = client.agents.create_agent(
    model="gpt-4o-mini",  # Using the low-cost deployment for data fetching tasks
    name="Data-Extraction-SubAgent",
    instructions=(
        "You are a specialized file path extraction agent. Your only job is to use the "
        "provided Azure Function MCP tool to fetch structured JSON file paths from Blob and SharePoint. "
        "Always return the clean JSON output back to the Orchestrator."
    ),
    tools=[{"type": "mcp", "id": mcp_tool_id}]
)

print(f"Sub-Agent Created Successfully. ID: {data_extraction_agent.id}")


Step 3: Bundle the Sub-Agent into a Toolbox (A2A Setup)

To allow a primary Orchestrator Agent to call this sub-agent, the sub-agent must be wrapped into an Agent-to-Agent (A2A) tool and bundled inside a Foundry Toolbox.

python

# 4. Wrap the Sub-Agent as an A2A Tool definition
a2a_tool = {
    "type": "agent_to_agent",
    "settings": {
        "agent_id": data_extraction_agent.id,
        "description": (
            "Use this tool when the user requests file paths, file listings, "
            "or data ingestion tasks from Azure Blob Storage or Microsoft SharePoint."
        )
    }
}

# 5. Create or Update a versioned Foundry Toolbox with this A2A Tool
toolbox = client.agents.create_toolbox(
    name="EnterpriseRoutingToolbox",
    description="Toolbox containing data extraction sub-agents and routing policies.",
    tools=[a2a_tool]
)

print(f"Toolbox published with A2A capabilities. URI: {toolbox.mcp_endpoint_uri}")


Step 4: Implement the Model Router and Orchestrator Agent

A Model Router evaluates incoming requests and forwards them to the appropriate model based on complexity. For multi-agent orchestration, we will set up the primary agent to use a dynamic routing profile: it boots on gpt-4o for high-level semantic planning but routes sub-tasks out to the toolbox efficiently.

python

from azure.ai.projects.models import ModelRouterConfig, ModelRouterRule

# 6. Define Model Routing Rules (JSON-schema style routing)
# If a prompt contains phrases explicitly asking for simple status checks or raw lists,
# the platform can automatically downgrade the primary LLM to save token costs.
routing_config = ModelRouterConfig(
    default_deployment="gpt-4o-heavy",
    rules=[
        ModelRouterRule(
            condition="prompt.contains('status check') or prompt.contains('just list')",
            target_deployment="gpt-4o-light"
        )
    ]
)

# 7. Create the Orchestrator Agent utilizing the Model Router and A2A Toolbox
orchestrator_agent = client.agents.create_agent(
    model_router=routing_config, # Dynamic cost optimization
    name="Primary-Orchestrator",
    instructions=(
        "You are the main coordinator. Analyze user prompts. If a user asks to scan, "
        "retrieve, or process files from storage systems, delegate the task completely to the "
        "Data-Extraction-SubAgent tool. Synthesize its output back to the user."
    ),
    toolbox_id=toolbox.id # Injecting the Toolbox containing the A2A sub-agent connection
)

print(f"Orchestrator configured with Model Router and A2A Toolbox. Ready for ingestion.")


Step 5: Execute the Multi-Agent Flow

When a user asks a complex question, the primary orchestrator processes the reasoning via gpt-4o, identifies that it needs storage paths, routes execution natively via A2A to the sub-agent, which subsequently runs the python codebase on your Azure Function App MCP Server.

python

# Create a user thread
thread = client.agents.create_thread()

# Submit a complex multi-layered request
client.agents.create_message(
    thread_id=thread.id,
    role="user",
    content="Find all CSV sales reports modified yesterday in SharePoint and give me a summary report."
)

# Run the Orchestrator
run = client.agents.create_run(thread_id=thread.id, agent_id=orchestrator_agent.id)

# Process the stream
while run.status in ["queued", "in_progress"]:
    run = client.agents.get_run(thread_id=thread.id, run_id=run.id)

# Print execution transcript
messages = client.agents.list_messages(thread_id=thread.id)
for msg in reversed(messages.data):
    print(f"[{msg.role.upper()}]: {msg.content[0].text.value}")

Use code with caution.


Step 6: Verify Telemetry in Application Insights

Because both agents are bound to the Azure AI Project client workspace, your Application Insights dashboard will capture a cascading hierarchy trace:

  1. Top-Level Span: Primary-Orchestrator execution via gpt-4o-heavy.
  2. Child Span (A2A Route): Event showing delegation to Data-Extraction-SubAgent.
  3. HTTP Webhook Dependency: Post request flowing straight into your MyFoundryMcpApp Azure Function endpoint.

To monitor this routing overhead and cost tracking in App Insights Logs, run this query:

kusto

customEvents
| where name in ("AgentExecution", "ToolCall", "ModelRouting")
| extend ModelUsed = tostring(customDimensions.model_deployment)
| project timestamp, name, ModelUsed, duration

NEXT: enforce custom Input/Output Guardrails at the Orchestrator level to block sensitive system paths from being sent down to the sub-agent, or we can configure a Shared Thread Memory pattern so both agents can access a history of files pulled previously.

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Post