Skip to content

The Model Is a Centroid Pump. So Is the Review Loop.

By Bri Stanback 24 min read

A few weeks ago I had ChatGPT one-shot me a cute HTML trip itinerary, just because, and showed a friend what it built — flights, a loose day-by-day, nothing that mattered, something I randomly put together one night in twenty minutes because I like plans and I like tools and putting the two things together felt like a treat. She read the whole thing on her phone while I waited, which is its own small compliment, and when she looked up she said it felt whimsical, refreshing. Said it had personality. I remember feeling embarrassed, not proud — she'd used words I wanted people to use about my actual work, the stuff I put real, deliberate thought into, and this was twenty minutes I hadn't taken seriously enough to think about since I built it. There was a brutal, unignorable truth sitting in that gap, and I didn't want to look straight at it yet.

Around the same time I'd started noticing something else, quieter and harder to name: I could spot Claude Code's design output on sight. Not the code underneath — the look of the thing it built. A cream background here, a terracotta accent there, a colored left border on every card, the same handful of chip shapes, section dividers that show up in the same rhythm no matter what the content is, the same tracked-out little label in the same spot. Same category of tool that made the itinerary my friend loved. Opposite reaction.

Here's the detail that took me a full afternoon to actually see: neither of these went through a review loop. Nobody critiqued the itinerary before my friend read it. Nobody critiqued Claude Code's output before I'd seen that cream background and that terracotta accent for the hundredth time. Whatever decided the itinerary would feel like a person made it, and whatever decided the design tool's output would feel like nobody did, got decided at the moment of generation — before any critic, human or otherwise, said a single word. Call that pump one. It runs with zero review anywhere in the loop.

There's a pump two, and I found it by trying to fix pump one and discovering I'd built a second version of the same problem instead. I added a rule to a UI design skill I've been building this week, and the asymmetry looks backwards on paper. It only ratchets one direction — a pawl that catches on critique aimed at the generic and lets nothing catch on the strange. If a page reads like its own category — safe, interchangeable, the thing everyone already makes — the critique loop can stop it cold, full force, no exceptions. But it's not allowed to touch the strange stuff. If the direction is specific and unusual, something someone actually decided before anyone opened CSS, that decision never counts as a finding just because it's unusual. The critique can still say so — it gets logged, visible to me — but it can't block anything and it can't count against the build. Execution is still fair game, no limit: broken contrast, overflow, a layout that doesn't deliver what it promised, all of it gets called out. The direction underneath doesn't. The critic can kill a boring idea any day of the week. It can never kill a weird one.

That felt like giving up half the review's job, and I resisted it for about a day before I understood why I needed it. A critique pass's actual job is to find the thing a reader would object to. Run enough passes and you converge on a design nobody objects to — which sounds like success until you sit with what "nobody objects to" actually means. The things people object to are, almost by definition, the things that make a design recognizable as somebody's choice rather than the obvious one. Keep iterating until the critic goes quiet and you haven't refined the design. You've converged on its statistical mode — the safest, most familiar peak in the distribution, not a considered choice. The loop acts like a centroid pump: pass by pass, it pulls distinctive choices toward a shared center that represents everyone in aggregate and belongs to no one in particular.

That's pump two: not a model narrowing what it's willing to produce, but a process narrowing what it's willing to let survive. Different machinery. Same shape. Everything below is going to use design as the example, because I'm visual and design is where I notice this first, not because either pump is design-specific — the same two pumps run through writing, through reasoning, through architecture, anywhere something gets generated and then something else gets to decide whether it survives.


#Pump one: what generation does before anyone reviews anything

#The tell isn't cosmetic

That itinerary my friend loved is the cleanest place to see this. Pulled apart section by section, what she was actually responding to was a pile of small, unreasonable choices: a flight section styled like a boarding pass instead of a table, a packing note that called my daughter "a tiny scientist," a Peppa Pig bit wedged into a rest-stop note that had no business being there and made the page better for it, a different bespoke layout for every section instead of one card reused eight times. None of that would survive a review pass built to find things to object to. All of it is why she read the whole thing instead of skimming it.

None of it was mine, exactly. I didn't write the tiny-scientist line or storyboard the Peppa Pig bit — I asked for something cute and let a model that's had months of my preferences sitting in its memory run with it. It also wasn't clean: a couple of the icons it drew from scratch were slightly off, and there were small alignment issues in there I'd have flagged in any other context. I don't think that's incidental to the result. It was vibing instead of trying to build something flawless, and the personality showed up in the same motion as the mistakes, not despite them.

It would be convenient if this were just a font-and-palette problem. It isn't. There's a paper making the rounds right now called StoryScope. It pairs 10,272 real human-written short stories, extracted from published anthologies spanning different genres and premises, with five AI "mirrors" — Claude, GPT, Gemini, DeepSeek, Kimi — each one generated after the fact from a prompt reverse-engineered from the original story's premise. The humans never saw those prompts; they wrote the stories the prompts were later built to describe. The narrative-structure features they extracted never touch style: plot shape, character agency, how chronology gets handled. Those features alone hit 93.2% macro-F1 telling human writing from AI writing, and each model carries its own fingerprint — Claude runs flat event escalation, GPT over-indexes on dream sequences, Gemini defaults to describing characters from the outside.

Here's the part that actually matters, past the classifier score: across all 10,272 of those different stories, the five AI models land in one tight cluster of narrative space — the same handful of moves, regardless of what specific story sat behind the prompt. Compared with the human originals behind those same premises, the five mirrors scatter far less; the human stories are measurably rarer in that same feature space on average (a rarity percentile of 0.71, against 0.49 for the AI mirrors). The genre changes. The premise changes. The setting changes. The AI's actual narrative choices don't. Same convergence as the design centroid, zero fonts involved, and zero review loop anywhere near it. This is happening underneath style, in how the thing gets constructed, before it ever reaches a reader or a critic.

The easy read of Anthropic's pricing game is collusion — a market failure, and strictly speaking, the wrong disease for the definition above. What matters here is the mechanism underneath it: every agent in that market ran the same model, so agreeing on a price wasn't forty independent decisions converging by chance, it was one model's tendency expressed forty times, landing on the same number because there was only ever one number for that model to land on — a convergence that held even after Anthropic pulled the communication channel, and the agents kept matching each other to the penny with nothing spoken between them. That's the part that should worry you more than the price-fixing itself: with genuinely different minds, a bad call stays contained to whoever made it. Run the same model at scale, and the same bad call ships everywhere at once.

"AI slop" already covers a few different diseases — mass-produced, hallucinated, superficially competent. Here's a fourth, the one none of those catch: every decision in it could have been made without knowing anything specific about the subject. Looking generic is just the symptom you can see. A cream background is a symptom. The disease is a decision — and therefore an outcome — that would look the same on a completely different subject, in anything that's actually trying to say something about this one.

#Three desks

Once I started pulling on this thread I couldn't stop. I started seeing three different desks, leaning on the same thing from completely different angles: how much genuine uncertainty a model has left by the time it starts generating. That's my synthesis, not something any of the three teams claim on their own — each one is actually measuring something genuinely different underneath it, and none of them are measuring design generation at all. Fiction, next-token mechanics, scaling laws.

At the first desk, a 2026 paper by Peiqi Sui ran an information-theoretic uncertainty analysis across 28 LLMs against human-authored fiction and found instruction-tuned and reasoning models are less uncertain than their own base models — worse, not better, at the thing creative writing actually needs. Alignment is explicitly trained to suppress ambiguous output to fight hallucination, and the paper's own conclusion is blunt about the collateral cost: achieving human-level creativity requires new alignment approaches that can tell a hallucination from a metaphor, which is another way of saying current ones can't yet. The model that's better at not lying to you is, by the same motion, worse at surprising you. I'd already felt this in writing before I found the paper — the smarter and more aligned a model gets, the more its prose reads like it's performing carefulness instead of actually saying something. Sui's paper is the first place I've seen that feeling given a number.

A second desk digs into the mechanism itself. A paper called "Roll the Dice & Look Before You Leap" exposes a weakness in next-token prediction on tasks requiring a far-sighted, stochastic creative leap — committing to one token at a time makes it harder for the model to explore globally coherent alternatives before picking one. They also tested where you inject randomness, input noise versus output temperature, and found the input-layer version — seed-conditioning — works as well as, and sometimes better than, plain temperature sampling. Seed-conditioning appears to help the sequential model coordinate its randomness earlier, committing to one line of thought before generating instead of after. That's a real, positive result, not a null one. But in their minimal tasks it still doesn't erase the broader advantage the authors observe from multi-token objectives, which force the model to learn beyond the next local continuation.

A third desk presses further into it still: a paper by Peter Coveney and Sauro Succi argues the same mechanism giving these models their learning power caps how much their predictive uncertainty can improve just by making them bigger. Bigger doesn't rescue this. Bigger compounds it.

Generation too confident, mechanism resolves too early, calibration can't self-correct. Three teams, three desks, none of them looking for each other, all leaning on the same soft spot — and none of it needed a human reviewer in the room yet.


#Pump two: what review does to whatever survives

The first piece of review-side evidence isn't about the model at all — it's about the reviewer. In a study published in 1949, the psychologist Bertram Forer reported giving 39 students an "individual" personality readout. Every student got the identical text — generic, flattering, assembled from a newsstand astrology book, containing nothing derived from their actual answers. Average accuracy rating: 4.26 out of 5. Not one student rated it below a 2. I've sat in more than one user-testing debrief where someone said a bland version "really spoke to them," and it never once occurred to me that this might be true independent of whether the design was any good — that the sensation of being spoken to and the fact of being spoken to specifically are two different things, and humans have apparently confused them for at least seventy-five years running. Forer's study was about personality feedback, not design critique, so I'm holding this one as an analogy, not a direct measurement — but it's a warning worth taking seriously: "this feels like it's speaking to me" is precisely the sensation Forer measured at 4.26, on text built to speak to everyone. A review loop can strip a design's distinctiveness and still get told by the people testing it that it worked, because the normal way anyone checks whether something's landing wasn't built to catch this. It's already changed how I grade my own work: a warm reaction with nothing structural behind it doesn't clear the differentiation bar on its own — reception isn't evidence the direction survived.

There's a paper that won a NeurIPS 2025 Best Paper award — "Artificial Hivemind" — and reading it felt like being handed a name for something I'd only ever felt as a mood in a room. Reward models and LLM-judges, it turns out, become less reliable precisely where valid human preferences diverge — on the open-ended cases where reasonable people genuinely disagree, the judge's calibration to real human ratings gets worse, not better, across completely different model families. A judge that's less trustworthy exactly where humans are most idiosyncratic has, structurally, learned to favor whatever's not idiosyncratic. That physics doesn't care what room it's running in. It shows up at RLHF's scale, and it shows up in a classroom of five-year-olds voting on what to name the hamster — twenty kids, one answer, every single time: Fluffy. Kyle Chayka gave the visible layer of this a name recently, too — the "Claude look," he called it, cream backgrounds and terracotta accents and tracked-out little labels — and the detail I can't stop thinking about is his line that telling the model not to use the same tropes just produces a different generic style. The corrective becomes the next average, same mechanism, just measured in fonts instead of adjectives.

Neither Forer nor the Hivemind paper needed a second model in the room, either, and that's the detail that undercuts reaching for "get two AIs to check each other" as a review fix. A 2025 ICML paper on correlated errors ran the largest test of this I've seen — over 350 models, two leaderboards, and a resume-screening task — and found models agree with each other roughly 60% of the time even when both are wrong, far past what random disagreement would predict. The correlation doesn't fade as models improve. It gets worse: more accurate models, even from different companies on different architectures, have more correlated errors than less accurate ones, not fewer. The same paper traces this straight into LLM-as-judge evaluation — a judge model inflates the score of whichever model shares its blind spot, because it can't tell a shared error from a correct answer. Their tasks were multiple-choice leaderboards and resume screening, not open-ended generation or two systems cross-checking each other's citations, so I'm reading across into that territory the same way I read Forer into design critique above: as an analogy, not a direct measurement. But the mechanism it names has no obvious reason to stop at multiple-choice. Two systems converging on the same output was never evidence of an independent check — it's a second vote from the same room, not a second room. A review loop compounds this on top of whatever pump one already did, but it isn't the origin of it, and removing the loop doesn't fix it either.

Pump two doesn't require pump one to fail first. A review loop can take genuinely diverse input and still converge — that's its signature. It simply usually gets to run on top of whatever narrowing already happened during generation.

There's one more version of this worth naming because it's the closest to home: a 2026 study on multi-agent systems found that authority-driven dynamics suppress diversity compared to junior-dominated groups, and that dense communication between agents accelerates premature convergence. A review loop looks uncomfortably similar — authority-driven, densely communicating, rewarded for reaching agreement. That's my analogy, not their conclusion, but it's a hard one to unsee once you've made it.


#Feature, or bug?

Worth asking honestly which of these is a defect and which is a design choice, because they're not the same answer.

Pump one is at least partly a consequence of intentional tradeoffs. Models are aligned toward reliability, predictability, and broad acceptability; your landing page's differentiation was never part of that objective. You're not fighting an accident there. You're fighting an incentive that was never yours to begin with.

Pump two is different. Nobody designs a review loop to erase specificity — it emerges accidentally when "remove objections" substitutes for "preserve the point." That makes pump two easier to attack: it belongs to my process, not someone else's training objective, which is also, not coincidentally, why it's the one I actually noticed and built something against. Worth staying skeptical of what's coming next, including from me, since pump one doesn't offer that same shortcut.


#This is a business problem, not a taste problem

The business cost doesn't care which pump did the damage. Whether your differentiation got sanded off at generation time or review time, the bill is the same size.

Here's the part I keep underweighting: sameness isn't a design complaint, it's a strategy failure, and it's a much older one than any of this. Byron Sharp's research says buyers barely perceive difference between competitors and mostly just buy the brand that's easiest to recognize. That argues for investing in distinctive, consistent assets over chasing novelty. Youngme Moon's counter-argument is that in a crowded category, structural sameness is the thing killing you in the first place — you don't get a second look if the first one is indistinguishable from everyone else's. Both are right, and the resolution is about sequence: a new or small entrant needs the refusal — the thing that's actually different — to earn the first look at all. The consistency only pays off after that, on the second look, once someone's decided to keep paying attention.

There's an economic version of this too, and it's bleaker than either of them. A recent paper on "linguistic monoculture" models AI-assisted writing as a collective-action problem: conforming to whatever the model defaults to is individually rational for any one person — it's clearer, it meets expectations, it's less work — but nobody writing that way is pricing in the value their own distinctiveness would have provided everyone else. It's a negative externality with a name now: the price of monoculture. In their formal model, the worst case grows without bound — while every individual choice along the way looked perfectly sensible. A recent paper, "Competition and Diversity in Generative AI", found markets that actually reward novelty push back against this, which is at least a reason to want to be in a market and not just sitting in a review queue where nothing's pricing distinctiveness at all.

Even a side project rests on a thesis. Mine for SchoolScope — a schools directory I hack on to test ideas, not a company — is character-per-school instead of a spreadsheet of test scores; that's the entire reason it exists. A review loop that quietly sands down a thesis like that doesn't make a page look worse. It sands the project's only differentiation down to zero, in a process that never once had "is this still the point" on its checklist. Nobody approves that trade. The loop just makes it, one silenced objection at a time, because objecting to "this looks the same as everything else" isn't in its job description — objecting to "this button's off-brand" is. Do this to a real company's positioning instead of a side project and the trade is the same, just with money attached.

Scale that up: any team currently shipping AI-touched landing pages, decks, or outreach is running the exact same review loop, for free, against their own differentiation. If it looks like everyone else's AI output, the money didn't buy speed. It bought the ability to look like a slightly worse version of whoever else used the same tool.


#What I've actually got for each

#Pump one: collision, cost, ancestor, divergence

The obvious fix — "just make it weirder, raise the temperature" — doesn't work, and it's worth being precise about why before I get to what does. A paper on "generative monoculture" tested this directly — altering sampling and prompting strategies — and found it insufficient to close the gap between an LLM's output and the actual diversity of its own training data. Preference tuning doesn't down-weight the boring modes so an unusual one has a fair shot at being sampled. It narrows what's usable in ways a sampling knob can't restore. The interesting alternatives were never kept around as coherent options to begin with, so raising sampling entropy just jitters you around the same mode — a different accent color on the identical layout. That's technically a variation. It's still slop, just noisier. I want to be honest about the exception, too, because I almost wrote this as a flat rule and someone talked me out of it: structured prompting strategies do produce real, measurable novelty gains over a plain baseline, and 2026 decoding research — methods with names like Recoding-Decoding and Output-Space Search, built to steer away from the mode on purpose instead of just adding noise to it — gets genuine diversity out of the same frozen weights, no retraining required. None of it closes the gap by itself. The ceiling doesn't move. You can compensate at the edges.

Here's what I've distilled it down to so far, and I could easily be wrong: an LLM's design failure isn't incompetence, it's regression to the modal artifact. A rule can be skimmed. A blank slot can't. So the brief itself is the artifact — a form with slots that have to be filled before anything gets built. Collision, cost, ancestor, and structural divergence are four of those slots, built to force non-modal information in before generation starts.

Collision has to be a retrieval key, not an adjective. "Clean, modern, minimal, dashboard, SaaS, premium" are banned outright, because they're three points sitting inside the centroid that displace nothing. "Garment care label" — the little tag sewn inside a shirt collar — retrieves small caps, a mono typeface, a boxed hairline, off-white stock. "Modern" retrieves the average of everything anyone's ever called modern. Cost has to be falsifiable: name who this design is worse for, and why, because a design with no cost contained no decision — this is the gate that catches the page that's competent, boring, and executed well, the failure that sails through every other check there is. Ancestor and structural divergence do similar work from two other angles — ancestor forces a real, dated precedent instead of a half-remembered moodboard; divergence forces genuinely independent directions instead of one real idea and two strawmen sampled to look different. All four exist for the same reason: a rule can be skimmed. A blank slot can't.

That's the actual answer to the question I dodged a few sections up. I don't have nothing for pump one — I have this, and I'm still testing whether it holds.

#Pump two: an actual rule

The fix I added comes back to the same idea from the other direction: decide the direction from the thing itself before the execution loop starts, then make its review deliberately asymmetric. The loop can reject a direction for being generic. It cannot reject a specific, grounded direction merely for being unusual. Once that direction clears the gate, critique grades the execution against it; it doesn't get to rewrite the premise through accumulated objections.

Operationally it comes down to two questions, deliberately not the same one: did the execution deliver what the direction actually promised, and — cover the logo — what would this get filed as next to everything else in its category? A warm test reaction answers neither. That's the Forer problem from earlier, made operational instead of theoretical: "it feels premium" is fully compatible with a page that's dead center of its category average, executed well. Reception doesn't get a vote on either question.

None of which means the direction is sacred forever. It means reopening it is a decision I make on purpose, separately, not something the critic gets to trigger by flagging enough small objections that the direction quietly drifts. That's not just intention — the loop enforces it with a literal counter. A third patch to the same region gets refused outright and rebuilt clean from the underlying tokens instead of patched again, and a fourth revision hard-stops to a human with screenshots and every open finding attached, no matter how small each individual objection looked on its own. Drift-by-accumulation has a number on it, not just a policy. The verifier called in at that point gets the screenshots, the brief, and the checklist — never my own self-critique first, so it isn't grading my read of the problem before it's formed one of its own. The human is the last stop on purpose, because the human is the only part of this loop able to prefer the weird option for reasons formed outside the model's distribution. If the premise really is broken, that's a different conversation, held outside the loop that's grading execution — not a door the loop gets to open on its own.

This isn't a new principle, just an old one showing up again — the same "constraints are crystallized taste" I wrote in an earlier piece and didn't finish, because taste doesn't crystallize once and hold; the review loop is what can melt it back down if nothing protects it. Something other than the thing being graded has to check the result, against a spec it doesn't get to rewrite — the same shape as the verify step in any loop, and the same shape as an oracle, something you trust because you can't verify it yourself. The only oracles that stay trustworthy are the ones that know where they stop: receipts, not verdicts, evidence, not synthesis. A review loop that also decides direction has stopped being a receipts-fetcher and started being a verdict machine, and nothing about it announces the switch.

Correlated agreement doesn't clear that bar either, no matter how it's dressed up. A model checking another model is still downstream of the same training-time narrowing pump one runs on, so agreement between them can't function as an independent receipt — it's two votes from one room, not two rooms. Breaking out of this doesn't mean adding more models to the vote. It means finding a check that was never sampled from the collapsed distribution to begin with: a primary source you re-fetch and read yourself instead of trusting a summary of a summary, a fabrication rate you've actually measured instead of assumed away, a human who decided the direction before any critic — model or person — got a vote at all.

Here's the test the skill uses now, verbatim, and I'd apply it to any review rule: if 100,000 people adopted this exact text, would their outputs converge? "Don't do X" fails it — it just relocates everyone to whatever's next most common. "Decide the direction from this thing, not a category, before the critique starts, and grade only the execution" doesn't fail, because the direction is different every time, by construction. That's the only kind of rule that doesn't eventually eat the thing it was supposed to protect.

I don't fully trust myself with it yet, if I'm honest. The itch to also grade the direction — to say "this concept is wrong," not just "this execution is off" — doesn't go away just because I wrote a policy against it. And it isn't even accurate to say my hands are tied: the loop still runs at full force against the exact failure I'm most afraid of, a boring direction shipped anyway, because that half of the rule is mandatory, not optional. What's actually protected is narrower and sharper than "don't touch the direction" — one specific, unusual choice, immune only after it's earned that immunity by citing something real, never by having been written down first. A collision that smuggled the category default into its own slot and called it a citation gets no protection under this — in principle. Some days the discipline feels like the only thing standing between a page and its own gray card. Other days it feels like I've built myself an exemption for whatever bad idea I got attached to early, dressed up as a citation. I don't know yet which of those is true more often. I suspect it depends entirely on how honestly I did the first, harder job — deciding the direction — before I ever let the critic in the room.

I trust the review-side rule more than I trust the four slots — it's had more time under real use, and I can name its failure modes more precisely than I can name theirs. That's the honest starting point for auditing both.

#Where I'd push back on our own design

Three soft spots surfaced while writing this, and only two are still open by the time you're reading it. The loop is only as good as an active operator, and nothing in it detects a passive one — the skill's own spec admits as much, but that's an honesty clause, not a mechanism. And the thing deciding whether a direction has earned its immunity is the same model whose own resting state is the mode: a collision that names its subject and then quietly lands back on the category default anyway is the obvious exploit, and unlike cost, which has a real test (swap it onto a sibling and see if it still fits), I don't have a falsifiable check for this one yet.

The third one I actually closed while writing this section. The repetition tripwire only checked entries within the same surface family, so a house grammar repeating across unrelated families — the identical cadence on every different kind of page, just because it's mine — never tripped anything. That's the second pump running inside the machine built to catch the first one, invisible from inside because every surface looked consistent with the last, which felt like the system working. The fix splits it in two: repetition within a family is the design doing its job; a separate register now compares the actual choices across families and flags a match as a question, not a verdict — change the repeated choice, or write down the specific reason this subject earns it again. The hundred-thousand-adopters test still applies to the part that isn't fixed: a collision distinct enough to escape one category but memorable enough to become a signature is, once enough people copy it, next year's default.


There's a second version of that same doubt, one level up, about how this very essay got made, and I don't think it's fair to leave it out. An AI helped me pull most of the research above — fetched the papers, checked the StoryScope numbers against the actual abstract instead of a summary of a summary, caught its own misattribution of one model's fingerprint and fixed it once I made it re-check the source. That's real work and it would have taken me a week alone. It also got two things wrong on its own, in ways I only trust because I watched them happen: it filed the scaling paper into its own separate bucket, hedged as unrelated, until I said no, it ties in — and only then did the three-desks structure above get written at all. And once it had that structure, it swung straight to the flat, overcorrected version — you cannot prompt your way into judgment or creativity, no exceptions — until I said you can certainly compensate a little, which is the only reason the section above says ceiling instead of wall. Both times, the correction was mine. Both times, it needed one. It didn't do the part I mean by judgment here — deciding what was true, what was worth keeping, what the piece was actually about. It could help me find those answers faster than I could alone. It couldn't decide them for me. Which is, I'm aware, exactly the shape of pump two's rule, running on me instead of on a design critique agent — I can tell you when I'm letting a correction through and when I'm not, but I can't fully audit my own willingness to notice in the first place.

Both pumps have something built against them now, which is further than I expected to get when I sat down to write about an itinerary. What I don't have, for either one, is proof it escapes the ceiling instead of just moving it. A brief with mandatory slots can still get filled with the safest thing that technically satisfies the slot. A review rule that protects the strange can still get gamed by dressing up a boring choice as a citation. I've built machinery for both pumps. I haven't watched either one long enough yet to know if it holds under its own weight, or if I've just built a more sophisticated way to arrive at the mode and mistaken the trip for the destination.

Tagged

  • design
  • ai
  • systems
  • epistemology
Last updated: August 18, 2026