2026-08-12 · 12 min read

FLUX 2: five predictions we are willing to be wrong about

What this is. Speculation, labelled as such. Nothing below is inside information and none of it describes a model that has been released. Each prediction comes with the condition that would prove it wrong, so this post can be marked rather than reinterpreted after the fact.

Predictions about unreleased models are usually written so they cannot lose. The trick is vagueness: say a model will be “significantly better at prompt adherence” and you are right no matter what ships. So here are five specific guesses, each with a falsifier.

One: text rendering stops being the headline

Legible text inside images was the demo that sold the previous generation. It is close to solved for short strings, and the remaining failures are long paragraphs and unusual scripts, which are not what marketing screenshots show. We expect the next release to lead on something else: editing precision, or holding a subject across multiple images.

Wrong if: the launch page leads with a text-in-image comparison.

Two: the interesting axis is editing, not generation

Generating a striking image from nothing is a solved commercial problem, and it is not what production work needs. Production needs the fourth version of an approved image with one element changed and everything else untouched. That is an editing problem, and the gap between a model that can regenerate and a model that can amend is where the real workflow value sits.

Wrong if: the release ships without a dedicated editing or in-painting path and positions itself purely on text-to-image quality.

Three: the weights split into tiers, and the open one is not the best one

The pattern is well established across the field: a permissive small model, a heavier model under a non-commercial or research licence, and the strongest variant available only through an API. It is a rational structure for a lab that needs revenue and reputation at once. We expect it to repeat.

TierWhat we expectWhy it exists
Small, permissiveRuns on consumer hardware, licence allows commercial useAdoption, community tooling, and the goodwill that comes with it
Large, restrictedBest local quality, licence limits commercial useKeeps the strongest local option from competing with the hosted one
Hosted onlyThe variant used in the launch comparisonsRevenue, and control over what the model is allowed to make

Wrong if: the strongest variant is released under a licence that permits commercial use without a separate agreement.

Four: consistency features arrive before resolution increases

Resolution is the easiest thing to advertise and among the least useful to add. Anyone shipping real work is upscaling anyway. The thing that is genuinely hard, and genuinely blocking, is producing twelve images that look like they came from one shoot. Expect reference conditioning, character locking, or palette control to be the marquee feature rather than a jump in native output size.

Wrong if: the primary announced improvement is native output resolution.

Five: inference cost per image does not fall much

Distilled and turbo variants have already taken the easy wins on step count. A larger model with better instruction following tends to cost more per image, not less, and the published prices at the frontier have been broadly flat rather than collapsing. We expect headline quality to rise and the price per image to stay in the same band.

Wrong if: a same-quality image costs less than half of what the previous generation costs at the same resolution, within three months of release.

How to evaluate it yourself

Whatever ships, the launch comparisons will be chosen to flatter it. A more useful test takes an afternoon and uses your own work.

  1. 01
    Freeze a prompt set
    Ten prompts from work you actually shipped, written down before testing.
  2. 02
    Run the incumbent
    Same seeds, same settings. This is the baseline you are comparing against.
  3. 03
    Run the new model
    Identical prompts. Resist the urge to tune them to suit it.
  4. 04
    Score blind
    Strip labels and have someone else pick. Attribution biases the scoring.
  5. 05
    Price the difference
    Cost per kept image, not per generated image.
An evaluation loop that survives contact with marketing material. The point of step one is that the prompts are fixed before you see the new model.

The last step is the one usually skipped. A model that wins on quality and loses on cost per kept image is not an upgrade for a team rendering at volume.

Why licences are the part we watch hardest

Our catalogue is open-weight, so the licence tier is not an abstract concern for us. A model whose strongest variant is hosted-only is a model we cannot offer on the terms we promise, which is that anything you standardise on here keeps running if you take it in-house. Prediction three is the one that decides whether the next FLUX shows up in our catalogue at all.

We will come back to this post when the model lands and mark it honestly, including the ones we get wrong.

MegaBrain Gateway

500+ models. One API. No markup.

Use in Claude Code, Cline, Cursor, or any coding agent.

Try MegaBrain free →

Newsletter

Stay in the loop

Get the latest model comparisons and guides — no spam, unsubscribe anytime.

← All posts