How to Choose the Right AI Model in Copilot Studio (2026 Guide)

Choosing the AI model in Copilot Studio looks like a small decision. It’s one dropdown on your agent’s Overview page. But that dropdown controls how well your agent reasons, how fast it replies, how many Copilot Credits it uses, and even where your data is processed.

I learned this firsthand last quarter. A logistics client asked me to build two agents in the same week:

  • A Contracts Q&A Agent for legal and procurement. It had to answer questions like “Can we end this contract early if the vendor misses two SLAs?”
  • A Password Help Agent for IT that walks employees through a self-service password reset.

Same company, same environment, completely different needs. The contracts agent needed careful, multi-step reasoning. The password agent needed to be fast and inexpensive, because hundreds of employees use it every Monday morning. They ended up on different models, and I’ll use both agents as examples throughout this guide.

In this guide, you’ll learn:

  • Where to change the primary AI model in Copilot Studio
  • What the General, Auto, and Deep categories and the release tags mean
  • Which admin settings control model access and data residency
  • How your model choice affects Copilot Credit costs
  • How to add deep reasoning or a different prompt model without switching the whole agent
  • How to test models side by side before you publish

Model names, settings, and rates in this guide were checked against Microsoft’s documentation as of September 2026.

Table of Contents:

Why the AI Model Matters in Copilot Studio

Every agent in Copilot Studio has a large language model (LLM) behind it. With generative orchestration turned on, this model decides which topic, tool, or knowledge source to use, fills in the inputs, and writes the reply. If you’re new to building agents, start with my guide on how to create an agent in Microsoft Copilot Studio and come back here.

So when you swap the model, you change the brain that makes those decisions. A new model might follow your instructions more closely. It might also format dates differently or call a tool too eagerly. That’s why I treat a model change like a design change: it gets tested and signed off, just like a new topic.

One quick note. Copilot Studio now runs agents on different harnesses (the runtime layer between your agent design and the model). This guide covers agents on the standard harness, which is where most topic-based and knowledge-based agents run today.

Quick Answer: Which Copilot Studio Model Should You Use?

If you just need a starting point, use this table. The rest of the guide explains the trade-offs.

If your agent…Start withWhy
Answers high-volume FAQs or runs simple actionsA General model, such as the default GPT-5.5 ChatFastest replies and lowest cost
Handles an unpredictable mix of simple and hard questionsAn Auto model (GPT-5 Auto is in preview, so test it first)Adjusts effort per turn
Analyzes contracts or policies, or troubleshoots across systemsA Deep model, such as Claude Opus 4.7, or a General model with deep reasoning on specific stepsBetter multi-step reasoning and citations
Runs a simple summary or extraction promptA Mini prompt model, such as GPT-4.1 miniBilled at the lowest (basic) rate
Must keep data in your geographic regionA GA model without the cross-geo tag in your regionAvoids cross-region processing

For production agents, stick to generally available (GA) models.

How to Change the AI Model in Copilot Studio

The primary model is the one your agent uses for orchestration and responses. Here’s how to change it:

  1. Sign in to Copilot Studio and open your agent.
  2. Go to the agent’s Overview page.
  3. Find the Model section.
  4. Open the dropdown and select the model you want.
  5. Ask a few real questions in the Test pane to check the new behavior.
  6. When you’re happy with the results, Publish the agent.
Change the AI Model in Copilot Studio

The Test pane uses the new model right away, but your users only get it after you publish. You can switch models at any time.

Each model in the dropdown shows small tags. One tells you what the model is good at (its category). The other tells you how mature it is (its release type). Reading these tags correctly is most of the job.

Copilot Studio Model Categories: General, Auto, and Deep

Copilot Studio groups models by what they’re built for. Here’s how I explain it to clients:

CategoryBuilt forSpeedCostGood fit
GeneralEveryday chat and light groundingFastestLowestFAQs, summaries, drafting, translation, simple actions
AutoMixed workloads; routes each turn dynamicallyVariesVariesHelpdesk and employee agents with unpredictable questions
DeepDeliberate, multi-step reasoning with toolsSlowestHighestContract and policy analysis, long documents, multi-system troubleshooting

General Models: Fast and Low-Cost

A General model is your workhorse. It answers FAQ-style questions, rewrites and summarizes text, and triggers simple actions quickly and at the lowest cost.

For the Password Help Agent, General was the obvious choice. The questions are short, the answers come from one SharePoint page, and the agent mostly calls a flow. A reasoning model would only add waiting time.

Auto Models: Effort That Adjusts per Turn

An Auto model decides how much effort each turn needs. A simple “hi” gets a quick reply, while a harder question gets more thought. It suits broad employee agents with a messy mix of questions. The trade-off is that latency and cost become less predictable.

Deep Models: Multi-Step Reasoning

A Deep model (often called a reasoning model) works through a problem before it answers. It breaks the task into steps and handles long documents with citations.

For the Contracts Q&A Agent, questions like “Which carriers allow price changes with less than 30 days’ notice?” require comparing clauses across several agreements. The Deep model was noticeably slower, but its answers were more complete and better cited.

Model Tags Explained: Default, GA, Preview, Experimental, and Retired

The second set of tags tells you how ready a model is for production. This is where makers most often get into trouble.

  • Default: The model new agents start with, usually the best-performing GA model. Microsoft upgrades it over time, and it’s the fallback if your selected model becomes unavailable.
  • Generally available (GA): Models with no release tag. Tested and suitable for production, though some have regional limits.
  • Preview: Likely to become GA later, but not recommended for production yet.
  • Experimental: For trying things out only. Expect variable quality, latency, and possible timeouts.
  • Early access environment: Experimental models that appear only in environments on Microsoft’s early release cycle.
  • Retired: A previous default model after an upgrade. You can keep using it for up to one month.
  • Cross-geo: Your data might be processed and stored outside your organization’s geographic region.

One billing warning: if you publish an agent that uses a preview or experimental model, usage is billed at normal rates. “Preview” doesn’t mean “free.”

Copilot Studio Models Available in September 2026

Here’s what Microsoft documents for standard-harness agents in US commercial environments:

ModelProviderCategoryStatus
GPT-5.5 ChatOpenAIGeneralDefault (GA)
GPT-4.1OpenAIGeneralGA
GPT-5 ChatOpenAIGeneralGA
Claude Sonnet 4.6AnthropicGeneralGA
Claude Opus 4.6AnthropicDeepGA
Claude Opus 4.7AnthropicDeepGA
GPT-5 ReasoningOpenAIDeepPreview
GPT-5 AutoOpenAIAutoPreview
Mistral Medium 3.5MistralGeneralExperimental (cross-geo)
GPT-5.3 Chat, GPT-5.4 Reasoning, GPT-5.5 ReasoningOpenAIGeneral / DeepExperimental, early access only (US)
Grok 4.1 Fast (Non-reasoning)xAIGeneralExperimental, early access only (US)
GPT-4o, Claude Sonnet 4.5OpenAI, AnthropicGeneralRetired

Government clouds (GCC, GCC High, and DoD) still use GPT-4o as the default.

This list changes often, so trust the tags in your own dropdown over any blog post, including this one.

Admin Settings, Data Residency, and External Models

This is the part that gets projects stuck in security review, so check it before you start testing.

Cross-Geo Models and Data Residency

A GA (cross-geo) tag means the model is production-ready, but your data might be processed outside your region. Which models carry the tag depends on where your environment lives:

  • Anthropic Claude models are GA in the United States but cross-geo in Asia, Europe, and the United Kingdom.
  • GPT-4.1 and GPT-5 Chat are GA in the United States, Europe, and the UK, and cross-geo in Asia.
  • Mistral Medium 3.5 is experimental and cross-geo in all regions.

For a European client with strict residency rules, the cross-geo tag alone can rule out a model. I check Microsoft’s region table before I test anything.

How to Turn On Preview and Experimental Models

If a maker says “I can’t see GPT-5 Reasoning,” an environment setting is usually off. In the Power Platform admin center:

  1. Turn on Preview and experimental AI models for the environment. Makers can’t see preview or experimental models without it.
  2. For experimental models, also turn on Move data across regions. This is managed at the tenant level and allows data to be processed outside your region.
How to Turn On Preview and Experimental Models

How to Enable Claude, Mistral, and Grok Models

External models from Anthropic, Mistral, and xAI appear in the same Model dropdown, grouped under the provider name. An admin has to complete two steps first:

  1. Turn on external models for the environment (or environment group) in the Power Platform admin center.
  2. Allow each model provider in the Microsoft 365 admin center, for example Connect to Anthropic LLM, Connect to Mistral AI Models, or Connect to xAI.

These settings are separate from the preview settings, so turning on one doesn’t turn on the other. If an admin later removes access to an external model, agents using it switch to a suitable internal model, or show an error if none fits.

Microsoft also carries a strong warning for Grok 4.1 Fast (Non-reasoning). In Microsoft’s safety evaluations it scored lower on safety and jailbreak benchmarks and is more likely to produce harmful or explicit content. For anything customer-facing, I wouldn’t use it.

If you’re still getting your head around how environments, admin centers, and connectors fit together, my overview of what Power Platform is is a good primer.

How the AI Model Affects Copilot Credit Costs

Copilot Studio bills usage in Copilot Credits. The standard feature rates apply to every language model Copilot Studio provides. For example, a generative answer costs 2 Copilot Credits whether it comes from GPT-5.5 Chat or Claude Sonnet 4.6. Reasoning changes the math.

The Reasoning Premium

When an agent uses a reasoning-capable model, Copilot Studio bills two meters:

  1. The normal feature rate for what the agent did, such as 2 credits for a generative answer.
  2. Text and generative AI tools (premium) for the reasoning tokens, at 10 Copilot Credits per 1,000 tokens.

For the Password Help Agent, with around 600 conversations a week, that premium would add up fast. For the Contracts Q&A Agent, with maybe 40 questions a week from lawyers, it’s easy to justify.

So I always ask two questions: “How many conversations will this agent handle?” and “What does a wrong answer cost us?” High volume and low risk point to General. Low volume and high risk point to Deep.

When Microsoft 365 Copilot Licenses Cover the Cost

There’s an important exception. For employee-facing agents, usage is included (no Copilot Credit charge) when the user has a Microsoft 365 Copilot license and the agent runs under that user’s authenticated identity. If most of your employees are licensed, the credit math above mainly matters for unlicensed users, customer-facing agents, and autonomous agents.

To estimate costs before you build, try Microsoft’s Copilot Studio agent usage estimator.

How to Use Deep Reasoning Without Changing the Primary Model

There’s a middle path I use a lot. Instead of switching the whole agent to a Deep model, keep a fast primary model and turn on deep reasoning for specific steps.

  1. Make sure generative orchestration is on (Settings > Generative AI > Orchestration).
  2. In the agent’s Settings, turn on Deep reasoning (preview).
  3. In your agent instructions, add the word reason to each step that needs it.
Deep Reasoning Without Changing the Primary Model

The agent then uses the reasoning model where you asked, or where it decides deep thinking helps. Here’s the kind of instruction I used for the contracts agent:

1. Identify the supplier name and the question type (termination, pricing, liability, renewal).
2. Search the Contracts library for the matching agreement.
3. If the question compares two or more contracts, use reason to compare the relevant
   clauses and explain the differences in plain language.
4. Always cite the clause number and the document name.
5. If you cannot find the clause, say so and suggest contacting the legal team.

A few things to know before you rely on it:

  • Deep reasoning is still in preview and currently runs on the Azure OpenAI o3 model.
  • It’s available in the United States and the EU (excluding the UK) and makes no data residency commitments.
  • Every reason step slows the reply, and each one uses billable Copilot Credits. Use it on one or two steps, not every step.
  • Check the Activity page afterward. The activity map shows a separate deep reasoning node wherever it ran, and you can expand it to see the steps and data used.

Deep reasoning is also handy for agents that run without a user, which I cover in my post on creating autonomous agents in Copilot Studio.

How to Choose a Different Model for a Single Prompt

This is my favorite cost saver. A prompt tool (built in the prompt builder) can use its own model, separate from the agent’s primary model.

The Password Help Agent has a prompt that turns the user’s issue into a clean ticket description before handing off to IT. That job doesn’t need the main model. A small, inexpensive model handles it fine.

Here’s how to change the model for a prompt:

  1. Open the prompt in the prompt builder (for example, from the agent’s Tools page).
  2. Select Model at the top of the prompt builder.
  3. Pick a model from the dropdown.
  4. Select the three dots (…) and then Settings to adjust Temperature and other options.
  5. Test the prompt with real sample inputs, then save.
Choose a Different Model for a Single Prompt

The prompt model list includes a Mini category that you don’t see for the primary model. GPT-4.1 mini is the default prompt model. You’ll also find General options like GPT-4.1, GPT-5 chat, GPT-5.3 chat, and Claude Sonnet 4.6, plus Deep options like GPT-5 reasoning, GPT-5.2 reasoning, and Claude Opus 4.6.

Prompt Model Rates

Each prompt model is billed at a basic, standard, or premium rate:

Prompt rateCopilot Credits per 1,000 tokensCopilot Credits per 10 responsesTypical use
Basic0.11Mini models for summaries and simple extraction
Standard1.515General models for richer writing and document work
Premium10100Deep models for analysis and reasoning

That’s a 100x spread between basic and premium, so using a Mini model for simple summaries is an easy win.

Temperature and Other Prompt Settings

  • Temperature goes from 0 (predictable, the default) to 1 (creative). I keep it at 0 for anything factual.
  • The temperature slider is disabled for the GPT-5 reasoning model.
  • Microsoft still treats Anthropic Claude models in prompts as experimental, even though they don’t show a tag. They’re hosted outside Microsoft and aren’t recommended for production prompts yet.
  • Prompts in Copilot Studio use Copilot Credits, while prompts in Power Apps or Power Automate use AI Builder credits first.

If you’ve used prompts in canvas apps before, the idea is the same. My walkthrough of the Power Apps prompt shows the Power Apps side, and my guide to extracting invoice details with AI Builder and Power Automate shows prompts inside a flow.

How to Test and Compare Models in Copilot Studio

Never switch models based on one lucky answer in the Test pane. Here’s the process I use.

Step 1: Build a Realistic Test Set

Collect 20 to 50 real questions, including the odd ones people actually type. For the contracts agent, two paralegals wrote down questions they get every week. For the password agent, I pulled phrases from old helpdesk tickets.

Save them as a CSV file with Question and Expected response columns, in that order:

Question,Expected response
"Can we cancel the Harbor Freight Lines contract early?","Yes. Clause 14.2 allows termination with 60 days' written notice if two SLA breaches occur in a quarter."
"Which carriers can raise prices with less than 30 days' notice?","Only Swift Coastal Transport, under clause 7.4 of its 2025 agreement."
"What is the liability cap in the BlueRoute agreement?","The cap is 12 months of fees, as stated in clause 11.1."
"i forgot my password and im locked out","Explain the self-service reset steps and offer to raise a ticket if the account is locked."

A test set can hold up to 100 test cases, and each question can be up to 1,000 characters.

Step 2: Run an Evaluation

  1. Open the agent’s Evaluation page.
  2. Select New evaluation, and then select Single responses.
  3. Choose Import to upload your CSV. You can also write questions yourself or generate a question set from your knowledge sources.
  4. Under Select test methods, add Compare meaning (checks whether the answer means the same thing as your expected response) and General quality. If your agent calls flows or tools, add Tool use as well.
  5. Select Run.
Test and Compare Models in Copilot Studio

Step 3: Switch the Model and Run It Again

Change the primary model on the Overview page and run the same test set again. Comparing the two runs side by side makes regressions easy to spot. Results stay in Copilot Studio for 89 days, so export them to CSV if you need a record for sign-off.

Evaluation results showing scores for each test case in Copilot Studio

Step 4: Look Beyond the Score

A score is a starting point. I also check:

  • Latency: Lawyers will wait 15 seconds for a good answer. Someone locked out of their laptop won’t.
  • Formatting: Did the model follow my rules for dates, lists, and citations?
  • Tool use: Did it call the right flow or knowledge source, or did it guess?
  • Refusals: Did it politely decline out-of-scope questions?

For the contracts agent, the General model missed cross-contract comparisons and the Deep model got them right. For the password agent, both scored about the same, so the faster, cheaper General model won.

Pro tip: Keep a “golden” test set for every agent. When Microsoft upgrades the default model, rerun the same set within a day. If something breaks, turn on Continue using retired models (covered next) and fix your instructions during the grace period instead of in a panic.

What Happens When Microsoft Upgrades the Default Model

When a new default model arrives, the old one gets the Retired tag, and agents on the default move to the new model automatically. If you need more time:

  1. Open the agent’s Settings page.
  2. In the Model section, turn on Continue using retired models.
Continue using retired models toggle in agent Settings in Copilot Studio

You can then switch between the retired and upgraded models for 30 days after the upgrade. The preference also applies to future upgrades until you turn it off. I use this for regulated clients who must re-validate every change.

The Model Is Only Part of Answer Quality

When a client asks “Which model should we use?”, my short answer is: General for high-volume simple questions, Deep (or a single reason step) for analysis, GA only for production, and no cross-geo models if data must stay in your region.

But good knowledge sources and clear topics matter just as much. If your answers are weak, check your SharePoint list as a knowledge source setup, and consider adding knowledge files automatically with Power Automate so the model always has current content. For fixed, predictable steps like a reset request, a well-built custom topic in Copilot Studio or an agent flow is often more reliable than asking any model to improvise.

For bigger solutions, you can also split work across agents. A fast front-door agent can hand complex questions to a specialist agent running a Deep model. I explain the pattern in my guide to building a multi-agent solution in Copilot Studio. And if your team is still deciding between building in Copilot Studio or using what’s built into Microsoft 365, read my comparison of Microsoft 365 Copilot vs Copilot Studio first.

Best Practices for Choosing a Copilot Studio Model

  • Treat model changes as design changes. Rerun your test set every time you switch models, even between versions from the same provider.
  • Keep preview out of production. Preview and experimental models can time out, change behavior, or disappear, and published usage is still billed.
  • Check data residency first. Cross-geo and experimental models might process data outside your region, which can fail a compliance review.
  • Reserve reasoning for high-value questions. Reasoning models add a premium token charge on top of the normal feature rate.
  • Use the cheapest model that passes your tests. A Mini prompt model or a single reason step often beats switching the whole agent.
  • Know where the switches are. Missing models are usually blocked by the preview, data movement, or external model settings in the admin centers.
  • Publish to apply. Test pane changes are instant, but users only get the new model after you publish.

Frequently Asked Questions

What is the default AI model in Copilot Studio?

As of September 2026, GPT-5.5 Chat is the default primary model for standard agents in commercial regions. Government clouds (GCC, GCC High and DoD) still use GPT-4o. Microsoft upgrades the default over time, so check the Default tag in your model dropdown.

How do I change the AI model in Copilot Studio?

Open your agent, go to the Overview page, and select a new model in the Model section dropdown. Test it in the Test pane, then publish the agent. Users only get the new model after you publish.

Which AI model is best for Copilot Studio agents?

There’s no single best model. Use a General model like GPT-5.5 Chat for high-volume FAQs and simple actions, and a Deep model like Claude Opus 4.7 for contract, policy or multi-step analysis. Test both with the same questions before deciding.

Can I use Claude models in Copilot Studio?

Yes. Claude Sonnet 4.6, Claude Opus 4.6 and Claude Opus 4.7 appear under Anthropic in the model picker. An admin must turn on external models in the Power Platform admin center and allow Anthropic in the Microsoft 365 admin center first. Outside the US, these models are marked cross-geo.

Why can’t I see preview or experimental models in the dropdown?

The environment’s Preview and experimental AI models setting is probably turned off. Experimental models also need the Move data across regions setting. An admin manages both in the Power Platform admin center.

Does changing the AI model change my Copilot Credit usage?

Standard feature rates, such as 2 credits for a generative answer, apply to every model Copilot Studio provides. Reasoning models add a premium charge of 10 Copilot Credits per 1,000 reasoning tokens. Prompt tools are billed at basic, standard or premium rates depending on the model.

What is deep reasoning in Copilot Studio?

Deep reasoning is a preview setting that lets an agent use a reasoning model for complex steps while keeping its faster primary model for everything else. Turn it on in the agent’s Settings, then add the word “reason” to the instruction steps that need it. It requires generative orchestration.

Can one Copilot Studio agent use more than one AI model?

Yes. The agent has one primary model, but each prompt tool can use its own model, such as GPT-4.1 mini for simple summaries. Deep reasoning can also send specific steps to a reasoning model.

What happens when Microsoft retires the default model?

Agents on the default move to the new model automatically, and the old model gets the Retired tag. Turn on Continue using retired models in the agent’s Settings to keep switching between the old and new models for 30 days.

How do I compare two models fairly in Copilot Studio?

Run the same test set of real questions on the Evaluation page, switch the primary model, and run it again. Compare the scores, then check latency, formatting, tool use and refusals before you choose.

Conclusion

Choosing the right AI model in Copilot Studio comes down to matching the model to the job. Use fast General models for high-volume questions, Deep models or a single reason step for careful analysis, and a Mini prompt model wherever it saves money. Test with real questions, respect your region and admin settings, and treat every model switch as a change that needs sign-off.

Start with one agent this week: build a 20-question test set, run it on your current model, and then try one alternative. The results will tell you more than any comparison chart, including the ones in this guide.

You May Also Like

⏰ LIMITED-TIME OFFER

Join the SharePoint & Power Platform Developer Live Training

📅 Live training starts October 5, 2026
✓ SharePoint Development
✓ Power Apps & Power Automate
✓ Copilot Studio
🎁
FREE 1-Year Access to SPGuides.Academy

Enroll now and get access to all 9 academy courses at no extra cost.

Secure your seat before the batch fills up.
Get the live training plus the complete SPGuides.Academy learning library.

Enroll Now & Get Your FREE 1-Year Academy Access → View live training details and schedule
Power Apps functions free pdf

30 Power Apps Functions

This free guide walks you through the 30 most-used Power Apps functions with real business examples, exact syntax, and results you can see.

Live Webinar

SharePoint Integration Power Apps Form With Repeating Table [Invoice Management System]

Learn how to build an invoice management system using SharePoint integration and a repeating table.

📅 2nd September 2026 – 10:00 AM EST | 7:30 PM IST

Download User registration canvas app

DOWNLOAD USER REGISTRATION POWER APPS CANVAS APP

Download a fully functional Power Apps Canvas App (with Power Automate): User Registration App