Provider deprecation
A provider announces sunset for a model in active use. The harness picks a replacement inside the same family. The regression pass clears it before any production traffic moves.
c-84, sector 65, Noida
Clixlogix’s framework for choosing frontier, open weight, and specialist models across engineering, marketing, and evaluation work.
Most AI delivery runs one model across every job in the pipeline. The framework routes each job to the cheapest model that clears the bar.
Every job carries these four attributes. The model that runs it is the cheapest one that clears the bar on all four.
Most AI delivery today picks one model and uses it for everything. Ad copy, architecture calls, code review, extracting numbers from a regulated document. Same model, every job. The economics do not work on high volume, the quality does not hold on specialist work, and the compliance posture does not survive first contact with regulated data. One model as the answer to every question was always going to break somewhere.
Clixlogix treats model selection as a piece of delivery engineering, no different from picking the right database or the right queue. Every job that runs in a pipeline carries four things worth knowing about before a model touches it. How hard is the reasoning. What happens if the answer is wrong. What kind of data enters the prompt. And what the budget will bear at expected volume. The model that runs the job is the cheapest one that clears all four. That is the whole discipline.
Four things get asked before a model gets assigned to a job. Every job in the pipeline runs through the same four questions, and a model only gets the job if it clears all four. The name on the model does not count as one of the four.
What the framework is asking. How much thinking does the task actually need. What the task requires to come back correct on a real file.
Scoring anchorsScored 1 to 5, calibrated against the evaluation suite built for each engagement. A 5 is architecture design, threat modeling, or reasoning that has to hold across many steps on regulated data. A 3 is code review, structured extraction, or a campaign strategy brief. A 1 is an ad copy variant, an email subject line, or a classification call that only takes one turn.
What the framework is asking. If this answer goes to production and it is wrong, how bad is it.
Scoring anchorsScored on four levels. Critical covers code changes in a payment flow, a medical dosing calculation, or a regulated disclosure. High covers a customer email, an external investor update, or a financial report. Medium covers internal tooling and analytics dashboards where the audience can catch a bad answer. Low covers a brainstorm, an early draft, or a search over an internal wiki.
What the framework is asking. What sensitivity class does the data belong to, and which providers has the client already approved for that class.
Scoring anchorsScored on four levels. Regulated covers PHI, PII under GDPR or HIPAA, financial transaction records, and material information that has not yet gone public. Confidential covers signed client contracts, internal roadmaps, and unreleased product plans. Internal covers operational data and employee records outside the sensitive tier. Public covers marketing copy, published documentation, and open datasets.
What the framework is asking. At expected production volume, how much are we willing to spend for every 1,000 tasks the job runs.
Scoring anchorsSet at engagement start and reviewed every quarter. A tight budget covers content production at volume, ad variant generation, and bulk classification, where cost per 1,000 tasks is the whole economic story. A moderate budget covers code generation and structured extraction, where cost still matters and quality has room to move the number around. A loose budget covers architecture, judgment work, and low volume tasks with high strategic leverage, where the cost of a wrong answer runs orders of magnitude ahead of the cost of the model.
| Criterion | Scale | Anchor examples |
|---|---|---|
| How hard is the reasoning | 1 to 5 | 5: architecture design. 3: code review. 1: ad copy variants |
| What happens if wrong | Low, medium, high, critical | Critical: payment flows. High: customer emails. Medium: internal analytics. Low: brainstorms |
| What data enters the prompt | Public, internal, confidential, regulated | Regulated: PHI and financial. Confidential: contracts and roadmaps. Internal: operational data. Public: marketing copy |
| What the budget will bear | Cost per 1,000 tasks | Tight: content production. Moderate: code generation. Loose: architecture work |
Twelve job families. Six model families. Hover any marked cell for the route and the reason.
| Job family | Frontier reasoning | Fast frontier | Open weight, general | Open weight, fine tuned | Specialist | Judge and eval |
|---|---|---|---|---|---|---|
| Architecture and specificationEngineering | ||||||
| Code generation at volumeEngineering | ||||||
| Code review and critiqueEngineering | ||||||
| Structured data extractionEngineering | ||||||
| Evaluations and gradingEvaluation | ||||||
| Retrieval and searchEngineering | ||||||
| Campaign strategy and creative briefMarketing | ||||||
| Content and copy at volumeMarketing | ||||||
| Ad copy variants and email at scaleMarketing | ||||||
| SEO content and topical clustersMarketing | ||||||
| Image and video generationMarketing | ||||||
| Brand safety and moderationEvaluation |
Complexity 5, risk high, low volume, loose budget. Reasoning depth dominates the economics.
Complexity 2, risk medium, high volume, tight budget. Cost per 1,000 tasks dominates.
Complexity 4, risk high, moderate volume. Adversarial review benefits from a second model family checking the first.
Complexity 3, risk medium to high, moderate volume. Reasoning plus schema fidelity.
Complexity 3, risk high on regression signals, moderate volume. Consistency across runs dominates raw capability.
Complexity 2, risk medium, high volume. Retrieval quality determines the ceiling on every downstream answer.
Complexity 5, risk high on brand exposure, low volume. Judgment work with high leverage per output.
Complexity 2, risk medium, high volume, tight budget. Tuning enforces brand voice. Open weights hold cost at the required ceiling.
Complexity 1, risk low to medium, very high volume, tight budget. A/B testing closes the quality loop downstream.
Complexity 2, risk medium, high volume. Grounding requirement met by retrieval, not by frontier reasoning.
Complexity varies, risk medium to high, moderate volume. Visual models built for the purpose outperform general models on this class of work.
Complexity 2, risk critical on brand exposure, high volume. Auditable classifiers preferred over general model judgment.
The current primary model within each family is reviewed quarterly and logged in the delivery harness. The family assignment is the published position. The specific model is an implementation detail, updated as providers release and deprecate models.
The family assignments in the map above are the published position. The specific models below cover the current implementation across every family. Reviewed and updated quarterly. Last reviewed July 2026.
Judge and eval routing configures frontier reasoning models as evaluators against engagement specific rubrics. Providers vary by engagement.
Additional specialist classifiers for moderation, brand safety, and content ranking run per engagement and are documented in the routing configuration.
Choosing the family is the easier half. Making sure the right model actually runs each job every time, across every engagement, is where governance carries the load. Five mechanisms in the delivery harness enforce it.
The routing configuration holds the family assignment for every job. An engineer working on a task does not pick a model. The harness picks it, based on how the job was classified.
Regulated data cannot reach a family the client has not approved. The routing layer blocks it at the point the prompt gets built.
Every quarter, Clixlogix reviews each provider’s terms, data retention policy, training policy, and residency commitments. Any change gets assessed against every engagement that touches the provider.
Every exception gets a signoff and surfaces again at engagement renewal for review. Exceptions do not accumulate silently in the routing configuration.
A model swap moves to production only after the engagement’s evaluation suite clears every affected job classification on the client’s actual workload.
The framework is a moving position. Four triggers keep it current.
A provider announces sunset for a model in active use. The harness picks a replacement inside the same family. The regression pass clears it before any production traffic moves.
A provider changes pricing, or a new release shifts the cost per 1,000 tasks benchmark. Clixlogix rebenchmarks the tiers, and the routing rules follow the new economics.
An open weight release matches frontier quality on a job family. Where it clears the same quality bar at lower cost, the family assignment for that job moves to open weight.
Residency law changes, consent requirements change, or a provider updates its data handling policy. Routing rules update before the effective date.
A mid market lender piloted an AI workflow for commercial credit memos. The prototype looked good in demo: it read borrower packets, extracted financials, summarized the business, and drafted memo sections. The problem appeared in pilot files with policy exceptions.
Routine extraction held up. Document intake, financial spreading, and business summaries came back clean across the pilot set. Policy exception reasoning did not. The model softened a covenant exception, missed a guarantor caveat, and produced fluent but unsafe memo language. The failure was uneven, which made it harder to see: the same workflow that handled ninety percent of files correctly was the workflow producing memo language a credit committee could not defend. That failure triggered the eval and routing discipline described in this framework.
One frontier model ran every step, including high volume document intake and financial extraction where reasoning depth added nothing. The pilot's cost per file was set by the hardest task in the pipeline rather than by the work each step actually required.
Exception reasoning and memo drafting ran in the same call. A model asked to draft persuasive language while judging a covenant breach resolved the tension toward fluency. Softened language read as competent and passed casual review.
Policy answers came from model memory rather than from the lender's approved lending policy. Reviewers could not tell whether a threshold in the draft came from the policy document, the borrower packet, or the model's prior.
Quality checks were manual spot checks on whichever files an analyst happened to open. There was no repeatable signal, so a prompt change or a provider version change could not be shown to have improved or degraded anything.
One pilot file carried a covenant breach with a compensating guarantor structure. The assistant drafted polished memo language that described the breach as a timing variance and recommended proceeding. The draft never mentioned the escalation the credit policy required. An analyst caught it. That single file moved the engagement from prompt iteration to evaluation.
Clixlogix built a golden dataset from the pilot files, labeled the expected flags on each, and ran the prototype against it (Fig 1).

Extraction accuracy cleared the acceptance bar. Policy exception handling failed on a material share of exception files. Every failure shared the same shape. The model produced defensible sounding prose in place of an escalation. The eval suite turned that failure from anecdote into a countable metric, which is what made it fixable.
Clixlogix decomposed the workflow into job families before any model touched production. Each job carried its own score on the four criteria and its own evaluation set. Drafting split from decisioning. No single model call had to both judge a policy exception and write persuasively about it.
The evaluation suite ran continuously against every release candidate, comparing prototype failures against the routed workflow and flagging critical failures before any of them reached production (Fig 2).

Policy answers routed through retrieval against the approved lending policy, so every cited threshold traces back to a source document a reviewer can open (Fig 3).

Quality review moved from manual spot checks to a judge and eval pass, with analyst review on flagged escalations only.
The routing table below shows what the framework produced for this engagement (Fig 4).
| Credit workflow job | Original route | Production route | Why it changed |
|---|---|---|---|
| Document intake | Frontier model | Fast frontier | Classification and structure, not maximum reasoning |
| Financial extraction | Frontier model | Fast frontier with schema evals | High volume; eval catches missing fields and date errors |
| Policy lookup | Frontier model | Retrieval plus controlled generation | Must cite approved lending policy, not model memory |
| Policy exception reasoning | Frontier model | Frontier reasoning | Critical risk judgment across policy, thresholds, and borrower data |
| Memo drafting | Frontier model | Fast frontier or fine tuned | Drafting separated from credit decisioning |
| Quality review | Manual spot checks | Judge and eval plus analyst review | Regression signals needed to be repeatable |
| Regulated data routing | Prompt and app logic | Enforced routing policy | Data handling needed enforcement outside the prompt |

Every job family in the credit memo pipeline carries its family assignment, its data sensitivity class, its evaluation suite version, and the review history of every swap since production went live. A credit committee reviewer or an internal auditor can trace any single memo answer back to the routing rule that shaped it.
Routine extraction moved to lower cost routes. The price of a credit memo no longer tracks the hardest reasoning step in the pipeline. It tracks the actual work.
Policy sensitive reasoning stayed on stronger models, separated from drafting. Retrieval grounds every policy citation the model reaches for, on every draft.
Routes, evaluations, and model swaps became reviewable artifacts a credit committee can inspect (Fig 5).

A reviewer can open any memo and reconstruct the routing decision that produced it. Which family handled the reasoning. Why the framework picked that family. What the evaluation gate cleared before the answer reached the draft. The framework’s job is to make that reconstruction possible on every memo, not on the ones a committee happens to sample.
A production AI workflow fails the day the business cannot prove why an answer reached the workflow.
The thesis at the top of this page was that one model as the answer to every question was always going to break somewhere. The framework is what replaces that answer with a routing decision every senior technical reader can defend.
If you want to see what that routing decision looks like on your workload, the form below is the way to ask. Written pass, one engagement, no meeting required to start. Broader questions live in the section after the form.
The published model map is reviewed on a quarterly cycle. The current primary model within each family is reviewed continuously in the delivery harness. A quarterly review produces the version of the map published on this page. Between quarterly reviews, individual family assignments update only on trigger events described in Chapter 6.
Client-specific requirements go through the exception process from Principle 4 of governance. The requirement is recorded with signoff from the account owner and the technical lead, applied to the engagement routing configuration, and surfaces at engagement renewal for review. The published position does not change; the engagement runs on the client-approved configuration.
Every job family in the model map has a secondary family assignment configured in the delivery harness. During a provider outage, the affected job classifications route to the secondary assignment for the duration of the incident. Evaluation regression on the secondary assignment runs before initial deployment, so the fallback is production ready when the incident starts, not built during the incident.
Regulated data classifications restrict model family choice to providers with data handling terms signed by the client that cover the classification. In practice this excludes open weight hosting arrangements that lack the required contractual coverage, and limits the assignment to frontier reasoning, fast frontier, and specialist families operating under enterprise agreements. The specific providers vary by client agreement and are documented in the engagement compliance record.
Cost per 1,000 tasks is tracked per job family in the delivery harness. Clients receive a monthly cost report broken down by job family, model family, and volume. The report includes any provider pricing changes during the reporting period and the family level cost trend against the engagement baseline. Quarterly reviews cover cost drift, capability parity assessments, and any routing rule updates.
A new frontier release triggers evaluation against the current primary model in the affected family. Clixlogix runs the engagement evaluation suite on the new model across the job classifications that touch the family. Where the new model clears the quality bar at equal or lower cost, it becomes the candidate primary for the next quarterly review. Production traffic does not move until the swap process from Principle 5 completes.
Fine tuning is considered when brand voice, domain vocabulary, or structure specific to the task requires enforcement across high production volume, and where prompt engineering has plateaued on the evaluation suite. Prompting is the default for lower volume jobs, judgment heavy jobs, and jobs where the task specification changes frequently. A fine tuned open weight family often clears the cost budget on high volume jobs that would otherwise require frontier reasoning with elaborate prompting.
Every engagement has a custom evaluation suite calibrated to the client job classifications. The suite covers accuracy on gold standard examples, regression against production traffic samples, and cost benchmarking at expected volume. For regulated data engagements the suite adds evaluations specific to compliance. A model swap requires the new model to match or exceed the current model on every metric. A swap that improves cost but degrades quality does not clear the regression pass.
Yes. Every engagement has a routing configuration record and a model swap log accessible to the client technical lead. The record shows the model family assignment per job classification, the specific primary model, any exceptions applied, and every swap that ran with the evaluation regression evidence attached. Clients typically review the log during quarterly business reviews.
The framework holds its position at the family level, independent of provider. Every job family has an approved secondary provider within the same model family, and evaluation runs across candidate providers on a rolling basis. Clixlogix does not accept commercial arrangements that would restrict routing choices or introduce preference into the framework. Partnership status with NVIDIA, Google Cloud, OpenAI, and Microsoft covers technical enablement and does not extend to routing preference.