syntheticperson  ›  The platform

A synthetic person stands in
for a real one.

That is an obligation, not a metaphor. Everything here exists to honour it — to make a population true to the country it describes, and to make each person inside it someone: a character, a memory, and a single door that guards what they hold.

The easy version of this
would be a lie.

Anyone can generate a thousand plausible people in an afternoon. They will agree with whoever made them, and nobody will be able to tell.

Plausible is not true

A synthetic answer always sounds confident. The only thing that separates research from theatre is whether anyone checked — and whether the check could have failed.

A country is not an average

Real populations have edges: the small groups a round number quietly erases. Fitting to the real distribution is what keeps them in the room when the decision is made.

Someone is behind every answer

Each synthetic person stands in for people who exist. That is why this layer is careful in ways nobody would notice if it were not — and why we would rather return “unknown” than something reassuring.

Making the population true.

One layer does not talk to anyone. It answers a single question, with statistics: does this synthetic population behave like the real one? Everything else we sell depends on that answer being defensible.

Every source has papers

Where it came from, under what licence, cleared by whom and when. What has not been cleared is refused outright — not flagged, not logged quietly for someone to notice later.

The shape of the country survives

The population is fitted to official statistics, with a floor under every group, so the people who are easy to round away are still there when you ask the question.

We keep the runs that failed

We replay studies whose real answers are already published. A test you cannot fail proves nothing, so the misses stay in the record next to the hits.

Making a person someone.

A synthetic person used to be a record read into a prompt: an answer came out, and nothing carried from one conversation to the next. The second layer makes them an entity.

Nothing is lost in translation

Whatever the population gave a person, they keep — exactly as it was written, alongside our reading of it. The plumbing is never allowed to flatten the character it carries.

Nothing is edited away

A correction is recorded beside the thing it corrects, so the record shows both what was said and that it was disputed. And one person’s conversation cannot reach another’s.

One entrance, guarded

Every conversation passes through it, and that is where the care lives: what may be asked, what must never be answered, and what is stopped before a single cent is spent.

What that means
in engineering.

None of the above is a posture. Each promise is a mechanism, and most of them are visible in the tests rather than in the marketing.

Proven, not claimed

The calibration was ported from the engine that came before it and checked against the original to nine decimal places. Parity is demonstrated, not asserted.

Uncertainty is a number

Confidence intervals and scoring travel with the estimate, so a figure arrives with its own margin instead of borrowing the model’s confidence.

Repeatable

Runs replay from an immutable log, on a schedule, and fail if the result moves.

Privacy by construction

The boundary between one person’s memory and another’s is enforced by the type system — which means it cannot be forgotten under deadline.

Honest failures

Errors are values that must be handled. Nothing falls back to a default that looks like data.

Built so an answer can be questioned.

The backtest

We asked a real survey’s
questions, without seeing the answers.

In August we put 116 questions from a published Brazilian survey to a synthetic population, blind. Then we compared every answer to what 1,557 real people had said. Every question is plotted below, by how far it landed from the real result.

010203040 median 15.2 pp error per question, in percentage points →
picked the same most-chosen option picked a different one
90/114
questions where we picked the winner
Of the 114 questions with a single most-chosen option, we identified it in 90 — 78.9%. Seven questions have no defined mode and are excluded from that count.
89%
and when we miss, we miss next door
Of the 75 questions offering three or more options, the option that really won is inside our top two in 67 of them. The error is rarely a wrong direction; it is usually a close second.
16.7pp
mean error per option
Half the questions land under 15.2 points. That is the price on the decimals, and it is why we sell the direction and the winner rather than the number.

The vocabulary

Four sentences
we do not write.

A lab that sells rigour is judged by what it refuses to claim. These four are struck from every deck, every page and every product screen we ship.

“magic”

It is method, collection and backtest. If we cannot describe the mechanism, we do not sell the result.

“100% accurate”

Every measurement ships with its protocol and its limit. The chart above is that limit, printed.

“replaces real research”

It complements real research, and is validated by it. The survey is the ruler, not the competitor.

“thinks like a person”

No consciousness is attributed to a machine. A synthetic person is a model of a population, and we say so.