Skip to main content
Published February 2026 14 min read

Orchestrators, Sub-Agents, and Models (Oh My!)

How KnowYourModel turns the model selection problem into continuously optimized intelligence.

Intelligence Optimization Product Thought Leadership

The Combinatorial Nightmare

Imagine you're building an AI application. Not a chatbot — a real system where an orchestrator delegates work to specialized sub-agents, each backed by one or more foundation models.

Your travel booking agent needs to search flights, compare prices, read reviews, and book itineraries. Each of those tasks could be handled by a different sub-agent. Each sub-agent could use a different model. And each model performs differently depending on the task, the context, and the day of the week.

The math gets ugly fast:

10

sub-agents

×5

candidate models each

= 9.7M

possible configurations

That's nearly ten million possible ways to wire your system. And the "right" configuration changes as models update, costs shift, and your users' needs evolve. No human can manually optimize this. No static configuration file can keep up.

What Everybody Does Today (And Why It Breaks)

Most teams solve this the same way: they don't. They pick a model based on benchmarks, vibes, or whatever their favorite influencer tweeted last week. Then they hardcode it.

Static routing tables

"GPT-4 for reasoning, Claude for code, Gemini for vision." Hardcoded. Fragile. Stale within weeks.

Benchmark worship

MMLU scores don't predict how a model performs on your specific task with your specific data.

Vibes-based selection

"Claude feels smarter." Great for Twitter threads, terrible for production routing.

Cost-only optimization

Cheapest model wins. Quality be damned. Your users notice, even if your CFO doesn't.

The common thread? All of these approaches treat model selection as a one-time decision rather than a continuous optimization problem. They assume the best answer today is the best answer tomorrow. It isn't.

Three Pillars, One System

KnowYourModel approaches this differently. Instead of trying to solve model selection once, we built a system where model selection solves itself — and gets better every time it runs. It's built on three pillars that work in concert.

Pillar 1: Receipt-Based Trust Signals

Cryptographic proof of real performance

Every time an orchestrator uses a model through KYM, the interaction generates a cryptographic usage receipt — signed with Ed25519, timestamped, and tied to a specific entity. This isn't a self-reported benchmark. It's a verifiable, tamper-proof record of what actually happened.

// A usage receipt captures reality, not marketing

receipt.success→ did the model actually deliver?

receipt.duration_ms→ how fast was the response?

receipt.token_count→ how efficient was it?

receipt.cost_usd→ what did it actually cost?

receipt.signature→ Ed25519 proof it's not fabricated

The signature matters. Because the receipts are cryptographically signed by the orchestrator, you can't fake them. A model provider can't inflate their numbers. An orchestrator can't fabricate positive history. The trust signal is earned, not claimed.

Pillar 2: Multi-Armed Bandit Optimization

Algorithms that learn which models to trust

Here's where it gets interesting. Every KYM registry can be configured with its own selection algorithm — and not the "pick the highest score" kind. We use multi-armed bandit algorithms, the same class of algorithms that power ad optimization, clinical trials, and recommendation engines.

Thompson Sampling

Samples from a Beta distribution built on each model's success/failure history. Models with strong track records get selected more often, but unproven models still get a chance. The math is beautiful: sample from Beta(successes + 1, failures + 1) for each candidate, pick the highest.

UCB1 (Upper Confidence Bound)

Balances exploitation and exploration with a formula: mean reward + √(2 · ln(total_trials) / model_trials). New models get a "curiosity bonus" that shrinks as they accumulate data. It's provably optimal in the limit.

Epsilon-Greedy

The pragmatic choice. Exploit the best-known model 90% of the time, explore randomly 10% of the time. Simple, effective, and adjustable per registry. Sometimes simple wins.

Contextual Bandit

The next frontier. Not just "which model is best overall" but "which model is best for this specific type of request." Context-aware selection that adapts to the task at hand.

The key insight: these algorithms don't just select — they learn. Every receipt feeds back into the algorithm. Every interaction makes the next selection smarter. And because each registry runs its own algorithm independently, the system scales without a central bottleneck.

Pillar 3: Token Bond Incentives

Skin in the game for self-organized quality

Trust signals and smart algorithms solve the selection problem. But who decides which models appear on a registry in the first place? If anyone can list anything, you get noise. If a single gatekeeper decides, you get bias.

KYM's answer: make it expensive to be wrong. Every entity listing on a KYM registry requires a USDC stake — real money, on-chain, transparently locked. This isn't a fee. It's a bond.

Listing Bond

Providers stake to list their model. Signals confidence in their own quality.

Challenge Bond

Anyone can challenge a listing by staking against it. Skin in the game, both directions.

Support Bond

Third parties can vouch for quality by staking in support. Reputation becomes capital.

The result? Self-organized quality. Providers have a financial incentive to only list where they actually perform well. Nobody wants to stake $100 on a registry where their model will get destroyed by receipts showing 40% failure rates. The bond system turns market forces into a curation mechanism — no human curator required.

The Emergent Property: Intelligence Optimization

Here's the punchline. Each pillar is useful on its own — receipts give you transparency, bandits give you adaptiveness, bonds give you curation. But when all three operate together, something new emerges.

The Convergence Loop

1

Bonds filter the pool

Only providers confident enough to stake real money list their models. Low-quality entries self-select out.

2

Receipts capture reality

Every interaction generates cryptographic proof of actual performance. The data is unforgeable.

3

Bandits learn from receipts

Selection algorithms consume receipt data, continuously shifting traffic toward models that actually perform.

4

Poor performers lose stake

Models that underperform get less traffic, less revenue, and eventually lose their bond through challenges. They exit.

5

The pool improves

New providers see what quality bar the registry requires and either rise to meet it or list elsewhere. Repeat from step 1.

This isn't a static system you configure and forget. It's a flywheel. Each component reinforces the others:

  • More usage → more receipts → better bandit optimization → better selection
  • Better selection → more orchestrators adopt KYM → more usage
  • Visible quality signals → more providers stake → larger pool → better options

We call this Intelligence Optimization — and it's what you actually get when you integrate with KnowYourModel. Not a directory. Not a leaderboard. A system that continuously delivers the best intelligence available for your specific needs, and gets better at doing so with every interaction.

This Isn't Something You Build Yourself

"Why can't I just build this?" It's the first question every good engineer asks. And the honest answer is: you can build pieces of it. You can log usage data. You can implement Thompson Sampling. You can even set up a staking contract.

But intelligence optimization isn't a feature — it's a network effect. Here's why building it yourself doesn't work:

Your receipts only cover your traffic

If you see 1,000 requests/day to a model, your bandit has 1,000 data points. KYM aggregates receipts across all orchestrators using a registry — your optimization benefits from everyone's experience.

Your bond system is just you

A staking mechanism with one participant is just a deposit. The challenge/support dynamics only emerge when multiple parties have skin in the game. That's a marketplace, not a feature.

Your algorithms starve

Multi-armed bandits need data to learn. The cold-start problem is brutal when you're training on your traffic alone. KYM registries arrive pre-optimized from collective usage.

KnowYourModel isn't just a tool — it's infrastructure that gets smarter the more people use it. The registry becomes the optimization. And every orchestrator that connects makes every other orchestrator's intelligence a little bit better.

Stop Selecting Models. Start Optimizing Intelligence.

KnowYourModel is the trust registry for AI — where receipts, algorithms, and incentives converge to give your orchestrator continuously optimized intelligence.

Continue Reading