Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Deploy the model behind an Azure Machine Learning (Azure ML) managed online endpoint, then decide what role it should play in Microsoft Foundry Agent Service: the agent’s conversational model or a prediction service the agent calls as a tool. Those are different integrations. Foundry’s documented bring-your-own-model (BYOM) connection expects an OpenAI-compatible chat-completions API behind an AI gateway; a task-specific Azure ML scoring endpoint does not become a compatible chat model just because it is reachable over REST.
Should the custom model be the agent’s model or a tool?
Start with what the model does, not where it is hosted. If it generates conversational responses, use the connected-model route only if its API matches Foundry’s documented requirements. If it returns a specialized result—such as a classification, score, ranking, or forecast—expose it as a tool and let the agent call it when appropriate.
| Decision | Connected as the agent’s model | Called as a tool |
|---|---|---|
| Purpose | Generate the agent’s conversational responses | Return a specialized prediction or task result |
| API contract | OpenAI-compatible chat completions through the documented gateway pattern | Can retain a task-specific request and response contract |
| Foundry integration | AI gateway and connected-model setup | OpenAPI, function, or MCP tool, or a call implemented in hosted agent code |
| Key check | Chat API compatibility, gateway routing, and authentication | Tool schema, endpoint authentication and reachability, and result handling |
Microsoft describes the BYOM gateway pattern as follows: “Foundry Agent Service allows you to connect and use models hosted behind your AI gateways such as Azure API Management or other non-Azure managed AI model gateways.” See Microsoft’s Bring Your Own Model to Foundry Agent Service guidance. The page describes third-party models brought to Foundry; that use of “BYOM” is distinct from Foundry Models offered by Azure.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →How do you deploy a custom model to an Azure ML online endpoint?
An online endpoint is the serving endpoint clients call. A deployment beneath it holds the resources that run inference, while the model’s scoring code or container determines how requests are processed and responses are returned. As Microsoft puts it, “A deployment is a set of resources required for hosting the model that does the actual inferencing.” See Deploy Machine Learning Models to Online Endpoints.
#1 Best Overall
- Define the serving contract. Specify the request fields the model accepts, the response it returns, and how it reports errors. The contract must suit the intended caller—an application, a Foundry tool, or a chat-compatible adapter—not merely expose a generic REST route.
- Prepare the model and runtime. Package the model artifact with its inference environment and scoring behavior. Azure ML deployments can use local or registered assets. Microsoft recommends registering model and environment assets for production reuse and traceability.
- Create a managed online endpoint. Choose an endpoint name that is unique within its Azure region and select the authentication mode. Azure ML documentation lists key-based, Azure ML token, and Microsoft Entra token authentication; it identifies Microsoft Entra token authentication as the most secure option for production managed online endpoints. Configure client permissions for the identity that will make calls.
- Add a deployment. Configure at least one deployment under the endpoint. Use Azure ML’s supported inference and scoring configuration or, when needed, a custom container that supplies a serving stack such as TensorFlow Serving, TorchServe, Triton, or another compatible server. Follow Microsoft’s custom-container deployment guidance for that route.
- Validate and operate it. Test the serving contract locally where practical, deploy to Azure, invoke the endpoint with an authenticated request, inspect logs, and monitor the deployment. Confirm that real callers send the expected schema and handle both successful responses and errors.
Microsoft’s deployment instructions cover Azure ML CLI v2 and Python SDK v2. The Azure Developer CLI also documents deployment to a Microsoft Foundry or Azure ML Studio online endpoint; see Deploy to a Microsoft Foundry or Azure Machine Learning studio online endpoint using the Azure Developer CLI if that workflow fits your project.
How do you connect the endpoint as the agent’s conversational model?
Use this route only when the custom model is intended to generate the agent’s conversational output and can meet the documented API contract. Foundry’s BYOM flow connects a model through an AI gateway and supports models that implement OpenAI-compatible chat completions. A normal Azure ML scoring endpoint may instead accept task-specific JSON and return a prediction, so do not assume direct compatibility.
Rank #2
- Check the API before configuring Foundry. Verify that the model’s interface supports the required chat-completions behavior. If it does not, put an adapter or gateway in front of the Azure ML endpoint to translate between the supported chat API and the model’s scoring contract.
- Expose the model through a gateway. Configure routing to the model or adapter and choose an authentication method supported by the gateway and Foundry connection. The BYOM guidance discusses API key and OAuth 2.0 options; required key headers can differ by topology, so use the documented expectations for the gateway you selected.
- Create the connected-model setup. In Foundry, follow the BYOM flow to create an admin-connected model connection with the gateway base URL and suitable authentication, add the model, and create a prompt agent using it. Use the current labels and steps in Microsoft’s BYOM instructions, since service interfaces can change.
- Test an agent turn end to end. Verify gateway routing, authentication, chat request and response behavior, and how errors surface to the agent. A successful Azure ML endpoint invocation alone does not verify the Foundry connection.
For gateway-based deployments, Microsoft describes Azure API Management as an option for controls such as load balancing, throttling or rate limiting, and governance. Choose those controls based on the requirements of your route rather than treating a gateway as an automatic compatibility layer.
How can a Foundry agent call an Azure ML prediction endpoint?
For a task-specific model, keep its prediction contract and make the operation available to the agent as a tool. Foundry documents custom tools based on functions, OpenAPI specifications, and MCP servers, as well as code-based hosted agents. See Microsoft’s What is Microsoft Foundry Agent Service? documentation.
Rank #3
- OpenAPI tool: Describe the callable operation and its input and output schema so the agent can use the endpoint as a defined action.
- Function or MCP tool: Expose the prediction operation through one of Foundry’s other documented custom-tool patterns.
- Hosted-agent code: Implement the endpoint call in the agent’s code when the application needs to manage orchestration or response handling directly.
In each case, the tool integration does not replace the agent’s underlying conversational model. The agent uses that model to decide whether to call the prediction service and to interpret its result. Keep credentials out of user-supplied tool arguments, grant the calling identity only the access it needs, and validate tool inputs and outputs before relying on them in downstream actions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What should you verify before putting the full route into use?
Test both components and the connection between them. A model that works when invoked directly can still fail when called by the agent because the identity, network path, or schema differs.
Rank #4
- Identity and authentication: Confirm the endpoint’s selected auth mode, the caller’s permissions, and any gateway key or OAuth requirements. Test with the identity and credential flow used by the deployed agent.
- Network reachability: Check that the gateway or agent can reach the endpoint under the selected network configuration.
- Request and response schemas: Validate required fields, types, response structure, and any translation performed by an adapter or tool definition.
- Failure behavior: Exercise rejected credentials, invalid input, endpoint errors, and timeouts; decide what the tool or adapter should return to the agent in each case.
- Operational controls: Review logs and monitoring for the Azure ML deployment and, if present, the gateway. Apply governance, rate limits, and routing controls appropriate to the service.
Microsoft also documents managed compute in Foundry, but that is a Foundry compute concept, not evidence that an arbitrary Azure ML scoring endpoint is a connected chat model. See Managed compute in Microsoft Foundry for that separate topic.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

