Choosing the model that powers an AI website builder is not a benchmark exercise. It's not about which model scores highest on a coding leaderboard or which one is cheapest per million tokens. It's about a single, stubborn question: when a non-technical person types one sentence, does the model produce a website a designer would respect?
This post is the record of one such evaluation — the exact prompt, the site our current default model produced from it, and what the output tells us. The default will keep changing as new models arrive and existing ones improve; that's a moving target, and we'll keep re-evaluating. But the method stays the same: give the model a real prompt, show what comes out, and judge it openly.
The prompt — one sentence, no follow-ups
Here is the entire prompt. Nothing was added, clarified, or iterated. No reference images, no brand kit, no second message. This is what went in:
// Single prompt, no follow-ups, no reference images:
Create a striking, minimalist portfolio website for 'Elias Voss Photography' specializing in fine art landscapes and portraits. Monochrome + selective color aesthetic (deep blacks, crisp whites, and muted teal accents). Large, dramatic full-screen images with perfect typography overlay. Smooth gallery navigation, about page with artistic bio, print shop with elegant framing options, and journal. The site must feel premium, emotional, and gallery-like.
That's it. One paragraph. Six sentences describing a fictional fine art photographer — the aesthetic, the pages, the feeling. No wireframe, no colour hex codes, no font names, no layout instructions. The model had to infer all of it.
This is the test that matters, because it's the test that mirrors reality. Most people who use Ullbek don't write detailed briefs. They write something like this — a feeling, a vibe, a few nouns. The model has to do the rest.
The site it built — embedded live, right here
This is not a screenshot. This is the actual published site, loaded live in the frame below. Click around it — every page, every link, every gallery transition is real. It was built and published entirely by GLM 5.2 from the single prompt above.
If you're on a phone, the frame above is compact — tap Open ↗ to view the full site in a new tab. Either way, what you're seeing is the unedited output. No human touched the HTML, the CSS, the images, or the copy after the model finished.
What the model actually built — the structure
From that one paragraph, GLM 5.2 made a six-page site with a coherent information architecture, a consistent visual system, and original copy written in the voice of a real fine art photographer. Here's what it produced:
Full-screen hero
A fog-shrouded mountain landscape fills the viewport. Overlaid typography: "Where light becomes stillness." A scroll cue, a photographer's statement, and a curated selection of recent compositions.
Filterable archive
Thirteen plates in a masonry grid with category filters — Landscapes, Seascapes, Portraits, Forest, Desert. Full-screen lightbox navigation. Every image alt-tagged with descriptive, SEO-aware text.
Artistic biography
A first-person artist statement, a portrait of the photographer, and a chronological exhibitions list — solo shows, group exhibitions, a monograph, an award. Written as if by a real artist.
Print shop
Six limited-edition prints with edition numbers, prices in pounds, three size options each, and four museum framing choices — matte black, natural oak, gallery white, charcoal. Each with material descriptions.
Field notes
Five journal entries — "The hour before light, Lofoten," "On printing black," "One light, one direction." Written between frames, with dates, categories, and read times. A working journal, not a blog.
Enquiries
A quiet enquiry form for print reservations, portrait commissions, and studio visits. Plus social links — Instagram, Behance, Vimeo — and a newsletter prompt.
Notice what's not there: no placeholder text, no "Lorem ipsum," no broken links, no empty pages. The model invented a coherent artist — Elias Voss, working from the north coast since 2009 — and wrote every word of copy in that voice. The exhibitions are fictional but plausible. The print editions are numbered and priced consistently. The journal entries read like a real photographer's field notes.
Two sides of the same coin — the harness and the model
It's worth being precise about what this test actually proves, because it proves two things at once.
The model — in this case, GLM 5.2 — proved it can take a vague, emotional brief and produce a website with aesthetic judgement, not just code. That's the side most people focus on when they talk about AI website builders: is the model smart enough?
But the model alone can't build a website. It can write code. It can't see the result. It can't check whether the hero image actually fills the screen, whether the gallery layout breaks on a phone, whether the CSS has a broken rule, whether a link 404s. That's the harness — the agent layer that wraps the model. It gives the model the tools to write files, place images, verify the render, fix its own errors, and publish.
- The model has taste. GLM 5.2 made design decisions — palette restraint, typographic scale, negative space, voice — that the prompt never explicitly asked for. The prompt said "premium, emotional, gallery-like." The output was all three.
- The harness has discipline. The agent layer verified the render, caught errors, placed images at the right aspect ratios, and published — so the model's taste actually made it to a live URL, not just a code snippet.
- Together they produce a publishable site. A great model with a weak harness produces beautiful code that breaks in the browser. A great harness with a weak model produces a flawless render of something generic. You need both.
This is why we don't talk about the model in isolation. When you build a site on Ullbek, you're not just talking to an LLM. You're talking to a system — a model with taste, wrapped in an agent that can see and fix and ship. The Elias Voss site is proof that the combination works. The model wrote the code; the harness made sure it actually rendered, published, and stayed live.
Why we're showing our work
Most companies pick a model and don't tell you why. They might say "powered by AI" or name a model in a footnote. They won't show you the prompt or the output.
We're doing the opposite. The prompt is above, in full, in a code block you can copy. The site is embedded above, live, in a frame you can click around. The evaluation is visible.
Because this is the question that actually matters to anyone using an AI builder: not "is it AI?" but "is it good?" And the only honest answer is to show you.
So here's the invitation. Take the prompt above — copy it verbatim. Paste it into Ullbek and see what our current default model builds for you. The default will change over time — new models arrive, existing ones improve — but the method won't: real prompt, real output, shown openly.
Try the prompt yourself
Copy the prompt above, paste it into Ullbek, and see what GLM 5.2 builds for you. Your first dollar of credits is on us.