Optimization is NOT overconfident + How it works in probabilistic environments

6.10.2026
Toward the end of a very productive debate, I dispel the myth that Trevor Miles perceived of optimization as being overconfident, and I also go into detail on how optimization can be modified to properly work for decision-making under uncertainty.

I enjoyed my debate with Trevor Miles. It stimulated me intellectually, and it sharpened my thinking more than most conversations on this topic have. His article "Optimization Is Also Too Confident" answers my take on his System 1 / System 2 post, and it agrees with more of my argument than it disputes. The System 1 framing was his, and I liked it. What I added was the architecture: mathematics at the core of high-stakes decisions, AI around it.

Here is where we already agree: A mathematical model adheres to its specification - whatever you specify in it, from a constraint to a balance equation, it honors. So when a model is wrong, it is wrong against the world, never against its specification, and you debug it by finding where the specification departs from reality. An AI that invents an answer departs from the specification itself, and there is nothing to debug against. Reproducibility lets you rerun exactly what ran six months ago. And structural logic extrapolates: if a truck needs a driver, fifteen trucks need fifteen drivers, whatever the history contained.

Trevor's new claim is that optimization shares System 1's worst habit: overconfidence. I disagree, and I think the disagreement is specific enough to settle.

What overconfidence means (to me at least)

Overconfidence means presenting something as more probable than it is. When an AI hands me code and appends "This robust, production-ready solution handles all edge cases," that is overconfidence. It is a claim, and nothing backs it.

A solver makes no such claim. Trevor names overconfidence as the trait, and in the next sentence describes it as never saying how sure it is. He then writes that a solver "returns one plan, labelled optimal, with no probability attached." Those statements are inconsistent. Staying silent about probability is not the same as overstating it.

The solver also says more than Trevor allows. "Optimal" is a mathematical proof that no better solution exists for the question the model asked. If the solver stops at a time limit, it reports a bound and a gap: the plan is provably within that distance of the best possible. Numerical trouble shows up in the return status. I spent a long time in semidefinite programming, where this mattered daily: you get an answer plus a flag that the numerics look shaky. And if the model is infeasible, there is no answer at all. Ask for an x greater than 5 and smaller than 0, and the solver assigns no number to x. It returns the status "infeasible." Optimization does not always have an answer.

Every statement a solver makes is narrow, and every one is true.

Trevor's real complaint

What an optimized plan does not tell you: how likely it is to come true, and how badly it misses when it fails. Both questions matter. But there are questions about using the technology, not questions of the technology. In my world view, using the technology correctly is the modeler's responsibility.

Classical mixed-integer optimization comes from combinatorial decisions. The textbook case is a knapsack: items with values, a bag with limited capacity, choose the most valuable selection that fits. There are no probabilities anywhere in that problem.

You can use the same technology where probabilities matter, and I will show how below. But you cannot blame a technology for lacking a notion it was never built to have. The same holds for "what you see is all there is": a constraint nobody wrote does not exist, and completeness of the model is the modeler's job.

A technology that demands skilled modeling is not an overconfident technology.

Slack, fragility, and plans that jump

Trevor and I share an admiration for Nassim Taleb, and I enjoyed how much of this exchange ran on his vocabulary.

Trevor's right to say

"An optimal plan is, by construction, a plan with no slack."

and I consider that a feature. You told the model nothing about uncertainty, so it had no reason to keep any slack. He is also right about fragility:

"Take ten percent off one binding constraint and the plan does not degrade gracefully."

And he describes what planners call nervousness: a slightly different forecast makes the solver jump to a very different plan, though the two plans look almost equally good on paper. Inside an organization, every such jump ripples through the people who execute the plan and makes them more nervous still.

The asymmetry makes the cost concrete: unused slack costs an idle driver's wage, missing slack costs expedited freight and lost orders. The latter is arguably more expensive by a lot. The Decision Factory by John Brandon Elam and Adam DeJans Jr. lays this pattern out for logistics. I wrote a free summary of that amazing story here.

Sequential Decision Analytics to the rescue

All of these are problems you get when optimization runs in an uncertain environment and the modeler has not done their part. The fix is to modify the deterministic formulation so it fits non-deterministic inputs.

The basic building blocks of Sequential Decision Analytics: Decisions are made based on states and thus-far-revealed uncertainties.

This is the core of sequential decision analytics (Warren Powell's framework, which I build on). Here are the modifications, each with a parameter that gets tuned:

  • Do not offer 100% of a capacity to the optimizer. Multiply the actual available capacity by a factor below 1, so the optimizer treats the case as if it had less capacity. That factor is one tunable parameter.
  • Decide what the optimizer receives for each uncertain input. A point estimate is the default, and it is the wrong default. Meinolf Sellmann argues you should provide distributions, and he has built a solver that accepts them.*
  • The Decision Factory offers further options: pass a high percentile of the distribution instead of the expected value, for example the 80th or the 95th. Which percentile is again a tunable parameter.
  • Alternatively, pass two bounds that cover a large share of the probability mass instead of a single number.
  • Add penalties to the objective. One penalizes changing the plan between planning rounds. Another penalizes approaching a capacity bound, a softer disincentive than the capacity factor above. The exact parameterization of the penalty is, once again, tunable.
  • Combine the above approaches. A capacity factor with a percentile input and a change penalty is one candidate; another combination is a second candidate.

How do you choose among candidates of modified optimization formulations?

With an outer simulation as the thing that tunes for 'real-world performance' (which is simulated). Take realistic scenarios for the uncertain inputs (Trevor suggests working from a company's existing history, and that works exactly the same way here). Run each candidate modified optimization formulation through them. Every candidate is its own independent optimization model, prepared for uncertainty in a different way, and you instantiate each one explicitly. Take the percentile example with demand as the uncertain input. Most people enter the 50th percentile when they enter the expected value. Instead, instantiate candidates at the 50th, 80th, 85th, 90th, and 95th, simulate them all, and see which performs best in the simulated 'real world'. You might end up at the 90th, because the 80th and 85th leave too little slack and the 95th leaves too much. The same simulation tunes the capacity factor and the size of the change penalty.

By modifying a deterministic optimization problem so that it can deal with decision-making under uncertainty, it becomes a policy. We can create a whole bunch of (differently parameterized) policies, and simulate their real-world performance.

Trevor writes that the tuned parameters

"have to come from something outside the solver that compares its answers with what actually happened."

Agreed. The outer simulation is that something. It also produces the plan he calls robust, one

"not harmed much by the error it did not foresee,"

and it prices the slack an optimizer alone would call waste.

In contrast to stochastic optimization (which Trevor mentions, but is a technique that is utterly useless in reality), sequential decision analytics moves the scenarios out of the optimization model and into the simulation around it. Trevor argues that a solver pays its full cost at decision time, every time. In this architecture it does not. The expensive part is finding the right slack: simulating many candidate formulations over many scenarios. That happens offline. Online, the system solves one optimization problem, the formulation that won the simulation, with the right slack already built in. The online policy does not even have to be optimization-based, but optimization looks ahead in time natively, as long as the model describes the future properly.

Calibration and black swans

"A stochastic program's range is only as honest as the distribution you fed it." He is right: the distribution is another input someone chose. I do not consider that fatal. The machine-learning forecasting built over the past two decades is good, occasionally wrong, and far from useless. Conformal prediction, which Trevor recommends, fits this loop well: measure how far past predictions missed, and size the bands from that record. It should become best practice. Kudos for raising it.

Then there are true black swans, events nobody saw coming. If one hits, you are out of luck, and that holds for every planning method. Trevor's own words apply to his approach as well: a learned policy is "strongest inside the range it was trained on." The only protection against the unforeseeable is to risk nothing, and then you have no business.

What sequential decision analytics offers is the infrastructure for adapting fast once new information arrives. Take the COVID pandemic as the example. (I doubt it was truly unforeseeable, but grant the premise.) A few weeks in, new patterns showed in the data (e.g. people spending their money in online shops). That information is rough at first, but you can already act on it: update the distributions, rerun the offline simulation, and deploy the formulation with the new slack. A few weeks later, more data has arrived, and you update again. The loop is built, and rerunning it is cheap. That is what matters when the unforeseen happens. If Trevor has a better answer for events nobody foresaw, I want to hear it.

How close our approaches are in the end

If I understand correctly, Trevor uses deterministic logic to generate training data for an ANN of his. That is a fair way to use a model. But the logic then lives in the generated training data, not in the policy. Once the policy goes online, a neural network answers, and nothing evaluates the logic anymore. Whether it assigns fifteen drivers to fifteen trucks depends on what its training covered and how it generalized. The correctness guarantee in the neural network is statistical.

In a mathematical model, the correctness guarantee is architectural. Trevor argues that the evaluation of a neural net takes less time than the solution of an optimization problem. In most situations that is correct, whether it matters depends on the application context. If one wanted to trade the correctness guarantee for faster evaluation, it is totally possible to replace an optimization model with a neural network in Sequential Decision Analytics, too (and Trevors training approach is clearly well suited for that).

Starting from 100% rigor, moving toward practicability

I personally am someone who comes from a rigorous math background and on the way to making decisions under uncertainty, I had to learn to give up optimality guarantees (e.g. if my simulation has not evaluated using the 87.5th percentile of an input distribution, my optimization taking a 90th percentile input may be inferior to that). I only gave up analytical rigor when absolutely required for practical purposes on my path and the place we end up in allows me to do this trade-off one more time: If I wanted to give up the architectural guarantee that my planning policy will not deviate from its specification in favor of shorter runtimes, I could.

Starting form data-only, moving toward practicability with more rigor (and the effort it entails)

Trevor seems to come more from a data-oriented side, where throwing lots and lots of data at the problem until it cracks is the general idea. Trevor went the route of embedding more rigorous modeling into his data generation process to be able to work with less but higher-quality data.

As I pointed out, the way I came here, I built a system that can very quickly adapt to suddenly very different distributions. Also, in terms of explainability, I'm very happy with the approach of 'rigorous analysis before masses of data'. Compared to Trevor, I am ignorant as to how easy these properties would be to replicate on his data-oriented side.

The final destination

Both of us have arrived at functionally the same thing: an approach that allows us to make decisions under uncertainty which squeezes the best out of the available data that we have. In short we have functionally created the next best thing after a crystal ball. After all of this analysis, I personally am surprised to arrive at this conclusion that we have essentially crossed the same river from two different river banks - and ended up 1mm close to each other. Plus the option provided by my sequential decision analytics framework to close the gap completely. What a wonderful place the internet could be if this was how debates would go.

Optimization is not overconfident. It makes no claim about probability, and it is honest about every claim it does make. Making it speak to uncertainty is a modeling job. It is a hard job, and a solvable one. Trevor Miles, thank you for a conversation I will keep thinking about.

‍

----------

* check out Meinolf Sellmann's talk at the DecisionCamp 2026, where he demonstrated for a hospital-scheduling application that simply considering probabilistic inputs as having a distribution gives better results than treating their expected value deterministically. The cool thing about his talk is that he shows what happens when a ridiculously wrong distribution is assumed: the results of working with a wrong distribution are STILL always better than had one assumed the expected value as deterministic.

‍

‍

More Posts

Making a Software Module Future-Proof: Refactoring Legacy C++ to Modern Value Semantics
Two bugs. Two full days lost. The answer was not the next patch, but a deliberately modernized core module. A case study.
1.9.2026
Success Story: Inventory Optimization Under Uncertainty at Dryft
When tight deadlines met complex uncertainty in inventory optimization, trust and innovation turned challenge into opportunity: a seven-figure cost reduction while improving service levels.
8.11.2025

Start improving your decisions today!

Unleash the power of modern software and mathematical precision for your business.
Start your project now