Generative Modeling Difficulty Hierarchy
Alex Mashrabov's map of what generative video can and cannot render, inherited from computer graphics: hard goods are easiest to model, liquids and spray are hardest, and virtual try-on scales with required precision. Glasses, headphones, and basic clothing work today, but premium fashion, which sells feeling and aspiration, needs Hollywood-grade, pixel-perfect fidelity the models cannot yet reach.
A physical-realism gradient that predicts what generative video can render convincingly today and what it cannot, carried over directly from computer graphics, where the same hierarchy has always held.
Explanation
Alex Mashrabov offers this as a scientific answer to which e-commerce categories generative AI can actually serve, ranked from easiest to hardest to model.1
Hard goods, rigid objects with fixed geometry, are the easiest to render convincingly, in both classical computer graphics and generative networks. Soft or structured wearables come next: virtual try-on of glasses, headphones, and basic shirts already works well today. Liquids and spray are, in his words, obviously the most difficult, since fluid dynamics are inherently hard to model convincingly. And premium or aspirational goods sit at the top of the hierarchy, hard not purely for physical reasons but because the fidelity bar required is so much higher: a Louis Vuitton dress, in Mashrabov's example, "sells not just the style, they sell certain feeling and aspiration... how it exactly should make a woman feel," which demands Hollywood-grade, pixel-perfect precision that current models cannot yet reach, making premium clothing try-on nowhere close to viable today.1
The key move in the hierarchy is that its hardest rung mixes a physical axis, the difficulty of rendering the underlying geometry and material, with a perceptual and brand axis, how close to perfect an image has to look before it can actually sell the product.1 Premium fashion is hard the way filmmaking is hard: not impossible to render, but unforgiving of the small imperfections that quietly destroy aspiration.
Why it matters
It functions as a practical go-to-market filter for generative-video companies: build for hard goods and basic apparel commerce now, and treat liquids and premium fashion as categories that are not yet served.1 It tells a company which e-commerce budgets are realistically addressable today rather than in some unspecified future. It also pairs naturally with a separate audience-acceptance axis, since the two together bound the addressable market on two dimensions at once, whether a model can render a category convincingly, and whether a given audience will accept the result as good enough. And the premium-as-filmmaking rung is, in effect, a taste-as-moat argument in disguise: wherever visual perfection is what actually sells the product, the residual human and precision premium survives the longest against automation.
Tensions and open questions
The hierarchy reflects the state of the technology as of the time it was described, and model progress could collapse individual rungs quickly, since basic virtual try-on already moved from hard to functionally solved within a relatively short period.1 It is a snapshot of a moving frontier rather than a fixed law. Conflating physical difficulty, such as rendering liquids, with perceptual difficulty, such as meeting an aspirational fashion bar, is a convenient simplification but mixes two genuinely distinct constraints; a model could in principle master fluid physics and still fail the aspiration bar for premium fashion, or the reverse.
Practiced by
Connections
Loading connections…
References
- 01
Alex Mashrabov on Higgsfield: Generative Video & Consumer AI (Venture with Grace)
Alex Mashrabov · interview · 2026
Related