Invariant LLM Systems

LLM solutions for production systems

Better decisions for every LLM task.

We are building Penelope, a cost-aware router that selects the right model and reasoning effort for each task. Alongside it, our review engine gives autonomous agents actionable feedback on code and research papers.

Flagship system
Penelope router
Review engine
Code & papers
Built for
Autonomous agents

What we are building

LLM infrastructure for better decisions and feedback.

Penelope is the core: it decides which model should do the work. The review engine then helps agents evaluate and improve what was produced.

01 · Flagship · Penelope router

Route every task to the model that should solve it.

Penelope selects one model and reasoning effort for every task. It needs no training: at inference time, it reads recorded evaluation evidence from similar work and makes one routing decision.

The goal is straightforward: preserve the quality that matters while avoiding expensive inference when a cheaper configuration is enough. Held-out masks, anonymised models, and explicit router cost keep the comparison honest.

Talk to us about routing
$40.13Model spend · lowest policy above 65 macro
38–55%Cheaper effort routing at matched accuracy
558kHeld-out evaluation trials in the evidence catalog
58 / 28Models / benchmarks represented

Honest total cost: the router currently adds $32.68. Penelope has the lowest execution spend, but not yet the lowest all-in spend.

02 · Review engine · code + papers

Give autonomous agents feedback they can act on.

The review engine generates a broad set of grounded concerns, removes semantic repeats, and selects a short list of specific, actionable feedback. It is a critic and improvement loop for autonomous agents.

On code, it reviews correctness, security, performance, and maintainability against the repository. On papers, it finds likely reviewer concerns while there is still time to revise the work.

Read the peer-review paper
84.9%Seriousness-weighted historical coverage
78.7%Strict issue coverage · ten-paper diagnostic
3,398Pre-review manuscripts in the study cohort
Top 32Target size for an actionable report

Broad generation works. Selecting the best short report is the open problem: current paper-only selectors retain 40–44% weighted coverage.

get in touch

Building an LLM product where cost or feedback quality matters?