The best multi-model workflow is not the one that calls the most models. It is the one that gives each call enough leverage to justify its cost.
That is the idea behind Conductor mode. GPT, Gemini, Kimi, and Claude are available in every new thread, but they do not have to repeat the same analysis. One model plans, the relevant specialists act, and the owner of each task reads the evidence it actually needs.
The published price gap
These representative standard API rates were checked on 6 September 2026. Prices are US dollars per million tokens; cached input, batch processing, priority service, tool charges, and gateway fees are excluded.
| Disidea role | Representative model | Input | Output |
|---|---|---|---|
| GPT | GPT-5.6 Sol | $4.00 | $20.00 |
| Gemini | Gemini 3.1 Pro Preview | $2.00 up to 200k prompt tokens | $12.00 up to 200k prompt tokens |
| Claude | Claude Opus 5 | $5.00 | $25.00 |
Gemini's published standard rate is the lowest in this representative comparison, GPT sits in the middle, and Claude is the premium call. Above 200,000 prompt tokens, Gemini 3.1 Pro lists $4 input and $18 output per million tokens, so the gap narrows for very large contexts.
This is a reference comparison, not a billing promise. The product uses moving gateway aliases, and the alias behind a seat can change independently of this article. Disidea therefore charges from the provider-reported cost of each real request instead of pretending every GPT, Gemini, or Claude turn has a permanent flat price.
For a rough 100,000-token input and 20,000-token output call at the short-context rates above, token cost would be about $0.80 for GPT, $0.44 for Gemini, and $1.00 for Claude. That is why putting a premium model on a high-leverage architecture decision can be rational while using it for a ceremonial final summary is not.
Kimi is a moving gateway alias rather than a fixed vendor build, so this article does not invent a permanent list price for it. Its real requests are billed from the cost reported by the configured gateway, under the same rule as the other moving aliases.
Four different strengths
| Model | Standing role in Disidea |
|---|---|
| GPT | Structures the problem, conducts the run, and owns general implementation |
| Gemini | Synthesises long context and reconciles competing views |
| Kimi | Develops practical alternatives and verifies implementation details |
| Claude | Challenges assumptions and handles architecture theory and structural review |
GPT is the general owner. It can turn a request into a bounded plan, implement ordinary product functionality, integrate work passed back from another model, and audit the result.
Gemini earns its place when the context is wide or conflicting. It is useful for tracing requirements across a long thread, reconciling alternative interpretations, and taking a bounded implementation handoff when a second active worker adds value.
Kimi adds another Advanced implementation perspective. It is available for practical alternatives, targeted verification, and bounded handoffs without taking over Claude's architecture-only role.
Claude is deliberately concentrated on structural decisions. For a substantial code request it reads the relevant repository files itself, records boundaries, contracts, risks, and acceptance criteria in an architecture document, and stops before implementation.
These are defaults, not claims that one model always wins. The panel is replaceable per thread, and an explicit model mention overrides the normal allocation.
Why there is no simple thread model
An earlier design kept a cheap model in every thread and routed repository reconnaissance through it. The savings looked attractive, but the handoff often added more uncertainty than value:
- The reader had to decide which files mattered without owning the downstream task.
- Its summary could omit a detail the architect or implementer needed.
- The next model still had to reopen important files to verify the summary.
- A weak tool-protocol translation could turn “I could not read the project” into the false claim that the project was empty.
The current rule is simpler: the model responsible for a task reads the files required for that task. The Conductor should not insert generic reconnaissance by default. If a genuinely separate inventory would prevent substantial duplicate work, the deployment can set a soft preferred reader with FILE_READER_MODEL; GPT is the default.
This removes an entire lossy handoff without removing inexpensive background processing.
Utility work is separate from the panel
Thread titles, Scribe briefs, and relationship maps do not need a frontier model seat. They share one deployment-level utility model selected with UTILITY_MODEL and UTILITY_PROVIDER.
The default is GPT-5 Mini. Operators can switch to Gemini 3.1 Flash-Lite, Kimi, Mistral Small 4, or GLM-5.3-Flash without changing the thread panel or application code. Provider-specific Scribe and title variables remain accepted only as migration compatibility.
This separation matters. A cheap background model should not participate in the argument merely because it names the thread afterward, and changing the Scribe should not change who is available to answer a user.
The Conductor pipeline
Ordinary questions stay ordinary. The conducting model may answer directly when a second view would add nothing. When several views are useful, it writes a visible one-to-five-step plan and assigns each step to the best available model.
Substantial project changes use two stages.
Stage 1: architecture, then implementation
- GPT writes the plan. It decides whether project files are genuinely required. The planning turn cannot modify the repository.
- Claude reads and designs. Claude opens the relevant files directly and writes one Markdown architecture document. It cannot write source code, workflow definitions, configuration, tests, or deployment files.
- GPT implements. GPT reads the architecture and the files it needs, then makes the real changes. During this implementation phase GPT and other selected non-Claude models may hand off bounded tasks to one another. Claude remains architecture-only.
No preliminary model summarizes the repository for Claude. No generic reader summarizes it for GPT. Each owner checks the source of truth.
Stage 2: audit, repair, and close
GPT first summarizes what Stage 1 actually did: which models ran, files changed, checks performed, failures, and unresolved problems. Free model handoff is disabled.
If problems remain, GPT repairs and verifies them directly. It then writes the final user-facing summary. A premium final narration would have little leverage after the architecture and code are already fixed.
Every model turn also has an explicit tool-call ceiling. Reaching it produces a visible progress summary with completed work and remaining work; the system never secretly removes tools and asks the model to continue as though nothing happened.
What this design improves
- Fewer repeated repository reads and fewer lossy summaries.
- Clear ownership: Claude owns architecture; GPT owns implementation and closure.
- Gemini and Kimi are enabled by default in every mode instead of being hidden in an optional pool.
- Utility-model cost can be tuned independently of thread quality.
- A provider or protocol failure is visible at the step where it occurs.
- Historical answers still render even when a model is retired from new threads.
What it does not guarantee
Claude can still choose the wrong boundary. GPT can still diverge from the architecture. Gemini can still over-compress a long discussion. A configured utility model can still return malformed JSON. Conductor makes these failures attributable and recoverable; it does not make them impossible.
The useful metric is therefore the finished task: tool success rate, retries, elapsed time, verified checks, total provider cost, and human correction required.
That is why Conductor is intentionally asymmetric. GPT, Gemini, Kimi, and Claude are not asked to contribute equal amounts. They are asked to own the part where their next token has the most value.
For the product's execution options, read Conductor, Sequential, Parallel — which mode and when.