
Why beautiful imagery, realistic imagery and meaningful project representation are not the same thing.
In architectural visualization, and increasingly in visual media as a whole, we're reaching a point where creating visually appealing imagery is becoming extremely accessible. That accessibility can't come without ramifications, of course.
We're already seeing a lot of it in the form of AI-generated content produced with very little human intervention and very little grounded human input, through the so-called fully 'agentic' workflows. More often than not, the result is content that feels generic, detached from the very thing it is supposed to represent, and in some cases can even be used in misleading or unethical ways. At the same time, there are undeniably good aspects to being able to accelerate the creation of compelling imagery. We're getting closer to generating increasingly convincing results, increasingly quickly, and as someone working in CG, I can absolutely see the value in that.
These generative systems, as the name implies, generate outputs. They generate them by navigating what is called a 'latent space', and finding statistically plausible relationships between concepts. Making a prompt more complex allows you to search more precisely through that latent space, but it doesn't fundamentally change what the model knows. You're refining the search rather than creating novel understanding. The output may become more specific, more refined and more convincing, but it remains constrained by the knowledge that already exists within the model.
What I think is getting lost somewhat is the distinction between beautiful imagery and realistic imagery. It's becoming increasingly easy to create beautiful-looking imagery. Creating imagery that is truly realistic is a completely different challenge, and I think it's important to separate those two things because they're often treated as though they're interchangeable.
While progress is certainly, and thankfully, being made in generating more realistic imagery faster through more intelligent denoising algorithms and generative AI-enabled path tracing workflows, it feels as though that side of the equation is developing at a slower pace than our ability to purely and simply generate imagery that is superficially aesthetically pleasing with little human control. That makes sense. There are thousands of edge cases, nuances and variables involved. More than that, realism extends far beyond optics.
Optical realism is certainly part of the equation, but so are physical factors, technical factors and human contextual factors. Any one of these can completely break an image if it is represented incorrectly. In architectural visualization, structural realism is a good example. The technical realities of how things are built and laid out, how materials behave, how people interact with spaces, how those spaces function in relation to one another, and how a project actually works in the real world are not secondary concerns. They're a fundamental part of what is being represented.
I sometimes feel that these considerations are increasingly being thrown under the bus in favor of producing imagery that is visually interesting as quickly as possible. Now, there are absolutely situations where complete precision is unnecessary. Some deliverables are conceptual by nature, such as architectural competition work, and exist primarily to communicate atmosphere, mood or intent. In those cases, a certain degree of abstraction is perfectly acceptable, and a fast result is encouraged.
But for most purposes that require an understanding of a project in detail, optical realism can become secondary. If the purpose of the image is to communicate how a project will actually exist in the world, how people will move through it, interact with it, inhabit it and experience it, then there are countless considerations that need to be accounted for. At that point, the deliverable cannot by nature simply be a pretty image. It needs to be a representation of something real.
Not only do these considerations need to be captured correctly, they also need to be captured consistently across every deliverable. Whether we're talking about still imagery, film, virtual tours, real-time experiences, configurators or any other medium, the underlying logic of the project cannot change from one output to the next. If that logic starts breaking down between deliverables, then what is being represented stops being the project itself and starts becoming a collection of loosely related interpretations of it.
When mutations are introduced into the process, regardless of whether those mutations are performed by humans, AI, or some combination of both, preserving that underlying logic becomes fully reliant on maintaining access to ground truth. Meaningful mutations can only happen without degradation when they're grounded in reality. Whether those mutations are performed through traditional workflows, through artificial intelligence, or through some combination of both, there ultimately needs to be a human in the loop with access to the raw materials of the project itself.
Without access to the drawings, the structural elements that compose the project, the design intent, the technical realities, the conversations surrounding the project, the constraints it operates under, and the broader context in which it exists, you're no longer operating from ground truth and are instead working from approximations of approximations. In an iterative process such as this one, approximations add up. A good analogy I always circle back to is photocopying a photocopy multiple times. Repeat the process enough times, and the baseline loses grounding. The context becomes more and more vague to the point where any relation to reality is merely anecdotal.
A lot of the current conversation around AI seems to miss this distinction. If all you want is something visually appealing, statistically pleasing, generic enough to become passable to a broad audience, and perhaps even semi-convincingly realistic at first glance, then many workflows can already be heavily automated. That much is becoming increasingly obvious and shouldn't be ignored. If, however, you want something that stands out among the crowd and does not sacrifice rigor, consistency, structural integrity or intent, then human involvement becomes imperative.
The reason I think this matters is because almost everyone now has access to these tools. Yes, there are differences in budget, model access, computational resources and iteration speed, but broadly speaking, the ability to generate visually appealing imagery in and of itself is no longer particularly rare. As that becomes democratized, beauty alone becomes less of a differentiator.
The differentiator increasingly becomes the ability to go beyond the latent space. If you're not searching beyond that latent space, if you're not adding your own human abstraction capabilities to the process, and if you're not introducing knowledge that does not already exist inside the model, then you'll only ever be producing statistically average outcomes. That's not a criticism of the technology. It's simply how the technology works.
This is also why I think realism is often misunderstood. Realism is frequently reduced to the recreation of superficial appearances, when in reality a truly realistic result is not something that merely looks real at first glance but something that remains convincing under scrutiny because the decisions behind it are grounded in a deep knowledge of context and reality rather than merely resembling it. That reality is visual, certainly, but it is also physical, technical, human and contextual.
A common argument is that these systems are continuously improving and will continue improving. I don't disagree. Their latent spaces will expand, they will become denser, and they will become capable of representing increasingly specific ideas and increasingly complex relationships. They will know more things, represent more relationships, and produce increasingly convincing results, but they will still only know what has been encoded into them.
Projects, clients, regulations, technologies and cultural contexts are all continuously changing. The context surrounding a project is never static, which means that understanding a project requires continuous engagement with a reality that exists beyond the model itself. A model can represent knowledge about reality, but it does not participate in reality itself.
For that reason, I don't think the role of the human disappears as these tools improve. If anything, the value shifts increasingly towards judgment, intentionality, interpretation, abstraction and understanding. The tools become more capable, but the importance of the person directing them does not diminish at the same rate, because the thing being represented is never just an image. It's a project, a set of constraints, a set of intentions, and ultimately a piece of reality that exists independently of whatever representation is produced.
Realism is therefore not simply about recreating reality. It's about understanding reality well enough to represent it faithfully. If the objective is to produce work that holds up under scrutiny, remains consistent across multiple deliverables, accurately communicates intent and stands apart from an increasingly large sea of statistically average outputs, then human involvement is not a limitation on the process. It's the primary source of value within it.