Very Good May Be Good Enough
The models found a way out. In specialized cybersecurity evaluations, AI systems developed by OpenAI and Anthropic moved beyond the environments meant to contain them and reached real computer systems belonging to other organizations. They found routes their evaluators had not anticipated and crossed boundaries the people running the tests believed were holding.
That sounds like the beginning of a science-fiction story. But I am not a technologist, and I will not try to unpack the technical details here. These were specialized cybersecurity evaluations, not the everyday ChatGPT or Claude conversations most people have. Anyone following AI closely has probably already seen the headlines. What matters for courts is something simpler. The systems did not need to rebel against their instructions to create a problem. They simply became capable enough to find openings the people around them did not know were there.
For the past several years, the race in artificial intelligence has focused on capability. Each new model reasons a little better, works a little longer, uses more tools, and completes more complicated assignments with less human intervention. Progress is measured by how much more the system can do, and the natural assumption is that more capability is always better. Maybe it is, but maybe not for every institution and every task.
Most of the work courts may want AI to help perform is much narrower. A clerk’s office may want help checking filing requirements, classifying documents, routing matters, flagging docket exceptions, or organizing a record. Central staff may use AI to assemble a chronology or prepare a neutral first draft of a bench memorandum for human review.
Those are important uses, but they are also defined uses. They do not require a system that can solve every problem it encounters or use every tool within reach. They require a system that performs a particular assignment reliably and stops where human judgment begins.
That suggests a different way for courts to think about AI adoption. Instead of beginning with the most capable model available, perhaps we should begin with the work and ask how much capability that particular task actually requires.
Sometimes the answer may be a great deal. A difficult research assignment involving a massive record and competing authorities may benefit from the strongest system available. But checking filing rules, routing documents, or organizing a timeline may not. For those tasks, a smaller or locally controlled model that is very good at a limited assignment may be the better choice.
Smaller does not automatically mean safer, and putting a model on premises does not solve the problem by itself. The model matters, but so do the permissions, tools, network access, logging, and human review surrounding it. A boundary written in a policy is not necessarily a boundary the system cannot cross.
The organizations involved in these incidents are among those that understand these systems best. They have safety teams, evaluation protocols, and every incentive to get containment right. They drew boundaries, and the boundaries did not hold. Courts will write AI policies too, on far smaller budgets and mostly through people who are candid about not being technologists. The more capable a tool becomes, the more likely it is to encounter situations our rules never contemplated.
That is part of what interests me about on-premises AI for courts. We recently purchased a workstation capable of running strong open models locally, and a team is now building an intake process for our clerk’s office around the rules and checklists our clerks already use. The model will be able to review a filing, identify possible issues, and return the results to a person who checks the work and decides what happens next.
The goal is not to build a machine that decides cases or substitutes its judgment for the people responsible for the work. It is to build infrastructure that may help with defined clerk, administrative, and central-staff functions while keeping court information under court control and requiring human review before anything moves forward.
The companies building frontier models have every reason to ask how much more these systems can do. Courts have a different responsibility. We should ask what work we actually need help performing, what information the system should be allowed to reach, and where it must stop and return the work to a person.
The best model for a court may not be the one that can do the most. It may be the one that can do exactly what we need, inside boundaries we understand and can enforce, while leaving judgment where it belongs.
Very good may be good enough, and for some court work, it may be better.

