Choose the right model
Evaluate whether a request needs a larger model or can stay on a smaller one. Reserve expensive calls for the work that needs them.
Request → modelSpecialized decision models · Early access
Train, evaluate, and deploy models for your business’s decisions.
Via the Zils Discord community. Customer delivery is planned.
CHOOSE THE RIGHT MODEL
context › Extract the total and due date from this invoice as JSON.
CONTEXT IN · PROBABILITIES OUT · EXAMPLES ARE ILLUSTRATIVE, NOT MODEL OUTPUT
A DECISION PRIMITIVE FOR YOUR AI STACK
Model selection. Tool choice. Escalation. Zils is building a trainable decision primitive for these repeated judgments: context in, probabilities out. The goal is to replace full LLM calls where a specialized model can meet your quality bar.
Potential savings depend on call volume, serving costs, and quality. Customer cost savings have not yet been measured.
Inside the AI stack
Train on authorized labeled examples from your workflows. Evaluate against your current stack, including quality, latency, and total cost.
Evaluate whether a request needs a larger model or can stay on a smaller one. Reserve expensive calls for the work that needs them.
Request → modelTurn context into a choice among the tools your agent can use, without generating a full text response for every selection.
Context → toolUse labeled outcomes to evaluate whether an agent should continue, retry, or ask for review before the next step.
Agent state → next actionProposed customer workflow
Customer-specific jobs and delivery are planned. Today’s research system already exercises training, checkpoint verification, and evaluation.
One recurring decision, its possible outcomes, and data you’re authorized to train on, with a separate set held out.
Adapt a starting model to the task. Recipes and candidate versions stay traceable.
Accuracy, probability quality, consequential mistakes, latency, and total cost vs. your current approach.
Acceptance criteria are agreed before training. Only a candidate that meets them moves on.
IT ALREADY RAN ONCE
This terminal replays the recorded Bittensor testnet round, line for line, from its public JSON. Every number on it is in the file.
testnet-round-001.json ↗Three miner processes and one validator were operated by one operator on one host. Fresh checkpoints; reused synthetic development benchmark, not an independent generalization test.
Planned deliverables
Know what was trained, how it was tested, and what it would take to put it to work.
A selected checkpoint with its identity, training configuration, and dependencies recorded.
Baseline comparison, dataset and rubric versions, measured tradeoffs, known limitations.
Self-hosting or a managed endpoint, assessed against your app, hardware, and data requirements.
Current research worker bundles include copies of training data; confidential distributed training is not supported. Data handling is part of scoping. No Zils model release is publicly available yet.
JEVBENCH · PUBLIC ITEMS
147/231
zils = published kev · +6 / −6 items
Open about the evidence
We publish the result as it came out, including where it got worse. This does not establish a general improvement.
Where do repeated decisions add up in your AI stack? Tell us about the calls, your labeled examples, and the quality bar a smaller model would need to meet.
Discuss your use caseOpens Discord. Start with a task description; don’t share private data.