Skip to content

Connecting Azure AI Foundry to Your Own Data

10 min read · updated August 11, 2026

A connection is a named, credentialed reference to something outside the project. Getting one working is mostly an identity problem wearing a data problem’s clothes.

What a connection is

Microsoft describes connections as the mechanism that lets a Foundry project authenticate to Microsoft and other resources, and as a prerequisite for building standard agents or using knowledge tools. In Azure terms it is a child object of the project holding a target endpoint, a credential and a type.

Two properties of that arrangement are worth holding on to. Connections belong to the project, not to the Foundry resource, so two projects on one resource can point at two different search services. And the secret does not live in the connection record: Microsoft documents that Foundry stores connection details in a managed Azure Key Vault, and that all Foundry projects use one. When you delete a project you are deleting the reference; the vault entry and the underlying resource have their own lifecycles.

The grounding target itself is a separate Azure resource you create and populate first. A connection to an Azure AI Search service with no index is a connection that resolves and returns nothing, and it looks exactly like a working one until you query it.

Two authentication modes, one of which is forced

A search connection can authenticate with an API key or keylessly with Microsoft Entra ID. The default in most walkthroughs is the key, because it works immediately.

There is one case where the choice is made for you, and it is the case most enterprises end up in. Microsoft states that if you use a private virtual network with the Azure AI Search tool you must use Entra ID project managed identity — key-based authentication is not supported with private virtual networking.

The practical implication is ordering. Build the connection keyless from the start even while the network is open, because retrofitting it later means re-testing every grounded path at the same moment you are also changing the network. Keyless needs the project’s managed identity to hold a data-plane role on the search service — reading an index is a distinct permission from managing the service, and the management role alone will not do it.

Create the connection

  1. Create and populate the Azure AI Search index. Confirm it returns results to a direct query before involving Foundry at all. Half of all “grounding is not working” reports are an empty or mis-analysed index.
  2. Enable a managed identity on the project if it does not have one.
  3. Assign that identity a data-plane read role on the search service, scoped to the service rather than the subscription.
  4. Add the connection in the project, giving it the search service’s endpoint and selecting Entra ID rather than key authentication. Name it something a flow or agent definition can reference without ambiguity — the name is an API surface.
  5. Test the connection from the project before wiring it into anything. A connection that resolves and a connection that returns documents are different tests; run both.
Role assignments take minutes to propagate, and a freshly created keyless connection failing immediately after setup is usually propagation rather than misconfiguration. Retry before you start changing things.

Grounding a query

With the connection in place, retrieval happens server-side. On the Azure OpenAI chat completions surface this is the data_sources extension: you pass the search endpoint, index name and authentication mode alongside the messages, and the service performs the retrieval, injects the results and returns an answer with citations attached to the message.

response = client.chat.completions.create(
    model="chat-default",
    messages=[{"role": "user", "content": "What is our refund window?"}],
    max_tokens=400,
    extra_body={
        "data_sources": [
            {
                "type": "azure_search",
                "parameters": {
                    "endpoint": "https://mg-search.search.windows.net",
                    "index_name": "policies",
                    "authentication": {"type": "system_assigned_managed_identity"},
                    "in_scope": True,
                    "strictness": 3,
                    "top_n_documents": 5,
                },
            }
        ]
    },
)

Three parameters do the tuning. in_scope restricts the model to the retrieved documents rather than letting it answer from its own knowledge. strictness sets the relevance threshold for including a retrieved chunk — higher means fewer, better-matched documents and more refusals. top_n_documents caps how many are injected, which is a direct control on prompt size and therefore on both latency and cost.

One interaction worth knowing before you hit it: the stream_options parameter is not accepted alongside data_sources, so a streaming client that was reading token usage from the final frame will start returning a validation error when grounding is added. That is discussed on the streaming page.

What retrieval actually sees

The model never sees your documents. It sees chunks — spans of text that the ingestion process cut out of them — and almost every grounding complaint that survives the parameter tuning above is a chunking complaint wearing a prompt’s clothes.

Chunk size is the trade. Small chunks retrieve precisely and arrive without the context that made them meaningful; large chunks carry their context and dilute the embedding, so the vector match degrades and each hit costs more prompt tokens. Overlap between adjacent chunks mitigates the first problem at the cost of duplication in results. There is no correct value, only a value tested against your own questions.

What breaks unrecoverably is a chunk boundary that falls in the wrong place. A table split across two chunks loses its header row from the second half, so the retrieved fragment contains numbers with no columns and the model will confidently misread them. The same happens to numbered procedures, to clauses that depend on a defining sentence three paragraphs earlier, and to anything where the meaning lives in the structure rather than the sentence. No strictness setting fixes a chunk that is wrong in isolation; the fix is in ingestion.

Two more properties of the index decide more than the query parameters do. Whether it holds a vector field at all — a keyword-only index will miss anything phrased differently from the source, and a query that works in your search explorer but not through grounding is often a vector field that was never populated. And whether semantic ranking is enabled, which reorders the top results by relevance rather than by raw score and typically matters more to answer quality than top_n_documents does.

Citations returned alongside the answer point at chunks, not documents. A user interface that shows them as document links is hiding the one piece of evidence that would let somebody tell you the retrieval was wrong. Surface the retrieved text.

When the answer is not grounded

Work down the chain rather than tuning the prompt, because the prompt is the last link and almost never the broken one.

  • Does the index return the document for the same query text, queried directly? If not, the problem is retrieval — analysers, chunking, or a vector field that was never populated.
  • Is the connection resolving under the right identity? A permission failure at retrieval time frequently surfaces as an ungrounded answer rather than an error, because the model happily answers from its own knowledge when it receives nothing.
  • Is in_scope set? Without it, an answer that ignores your documents is permitted behaviour, not a bug.
  • Is strictness too high? Raising it to suppress irrelevant results will eventually suppress the relevant ones. Move it one step at a time against a fixed question set.

Keep a small set of questions with known correct source documents and run it after every change to the index, the connection or the parameters. It is the only way to tell a retrieval regression from a model one, and building it takes an hour.