← Back to issue
a16z

Building the Future of Image Generation with Ideogram's CEO

41m · Transcribed via assemblyai · Watch on YouTube

Mohamad, founder/CEO of Toronto-based Ideogram (speaker A), on shipping their first **open-weights image model at 9.3B parameters — roughly 9× smaller than the ~80B 'SOTA'** and runnable on a single consumer GPU. The whole interview is a clean case study in a David-vs-Goliath wedge: **'we know we can't win on scaling'** — as an ex-Googler he's blunt that 'even if we raise 10x... we can beat Google in terms of the number of chips' is false — so the strategy is the inverse: win on *focus, differentiation and taste* in a niche the big labs ignore. Ideogram's wedge from day one was accurate **text rendering** (when image gen 'was synonymous with garbled text' and DALL·E 2 made meme travel-posters with wrong city names), which turned out to be 'the whole graphic design and storytelling industry.' Two ideas travel well beyond image models. First, **taste as a deliberate, measured moat**: 'we really want our models to have taste,' defined as 'going outside of the norm... not conforming to the average opinion — which is a little against being on top of the leaderboard'; he ran *very little RL* on purpose so the model stays stylistically diverse rather than converging on the same RL-flattened look every frontier model produces, and uses human designers (not AI) for side-by-side taste evals. Second, **the intermediate representation**: the model is trained only on JSON prompts (~4,000 tokens), so a language model expands a vague idea into structured JSON and the diffusion model renders it — 'making the task as straightforward as possible for the diffusion model' — and they *show users the actual model input* (unlike OpenAI/Google) to give control and consistency, likely migrating from custom JSON to HTML since LLMs already know it. Open weights is the GTM: partner with chip-makers, inference providers and enterprises who want on-prem, on-device and brand-DNA fine-tuning. And the build loop is already agentic — 'in a couple hours you have your landing page up and running' from an agent hitting their API/MCP. Direct line to last week's Fadell thesis: as models commoditise, taste and use-case fit, not raw capability, are the differentiator.

Key points

Notable quotes

It's not about how good a model is in the general sense. It's about how good is this model for my use case.

A · 0:03

We focused on the details of the model and we know we can't win on scaling.

A · 18:19

I used to work for Google. I don't think even if we raise 10x the amount we've raised so far, we can beat Google in terms of the number of chips that we can dedicate to each model training.

A · 18:31

One element of taste is kind of being— going outside of the norm a little bit and not conforming to the average opinion, which is a little against being on top of the leaderboard.

A · 16:37

We really want our models to have taste.

A · 0:22

I don't think a lot of labs are focusing on design, graphic design in particular, editable text that I'm talking about.

A · 18:51

you can go into your agents and then ask it to connect to the API and generate a bunch of images and then you can go and find the best ones and like in a couple hours you have your landing page up and running.

A · 30:17

the recipe for building more powerful models, in my opinion, is making the task as straightforward as possible for the diffusion model.

A · 36:48

Themes

Mentioned