Can an AI predict which synthetic faces humans think look intelligent?
Status: Public-report version 1.0
Date: 2 August 2026
Study status: Data collection and the frozen primary analysis are complete. This report was written after the primary result was known but before production of the planned video. It is not a peer-reviewed paper.
The most important sentence in this report: This study did not test whether intelligence or IQ can be read from a face.
Every face in the experiment was artificially generated. None of the depicted people exists, and none has an IQ, a biography, a personality, or any other psychological ground truth. The experiment tested whether one AI system could predict the average first-impression rating that human participants would assign to a synthetic face. Agreement between the AI and the participants is evidence of shared judgments about appearance. It is not evidence that either humans or AI can detect intelligence.
Abstract
Can an AI system predict which entirely fictional faces a group of humans will judge as looking more or less intelligent? To examine this narrow question, I created a fixed set of 72 synthetic adult portraits. The images shared a standardized portrait format but deliberately varied in visible appearance. One hundred and twenty Prolific participants each rated 18 unique faces, plus one concealed repeat, on a seven-point scale in response to the question: “Based only on immediate appearance, how intelligent does this fictional person look to you?” Each face received 30 independent non-repeat human ratings.
Separately, before the human results were inspected, a frozen AI-scoring protocol was run. The AI system saw each face in a fresh, isolated context and was asked to predict the arithmetic mean of the human ratings for that image—not to estimate actual intelligence. Five independent calls were made per face, and their mean was the face-level AI prediction. The prespecified primary analysis was the Spearman rank correlation across the 72 face-level human and AI means.
The observed Spearman correlation was 0.673. In 100,000 random permutations of the face labels, none produced an association at least as extreme; using the prespecified plus-one correction, the Monte Carlo p-value was 1/100,001, or approximately 0.000010. A participant-cluster bootstrap gave a 95% interval of 0.495 to 0.706. Pearson’s correlation was 0.665. The AI’s ordering was therefore strongly associated with the group’s ordering within this particular stimulus set.
The result is statistically clear for the frozen dataset, but its meaning is easy to overstate. It compares two highly aggregated scores, is conditional on a purposively varied set of synthetic faces from one generator and one visual style, and may largely reflect cultural stereotypes or other learned appearance associations shared by the participants and the model. The AI’s ratings were also lower and more compressed than the human ratings, so strong rank agreement did not imply close calibration. The result does not establish individual-level agreement, causal facial cues, generalization to real people, generalization to other populations or models, or any capacity to infer intelligence. It is best understood as evidence that an AI system can reproduce a substantial part of one group’s collective first-impression ranking of these synthetic portraits.
Plain-language summary
The experiment asked two related but very different questions:
- Which of 72 synthetic faces did a group of human participants tend to rate as looking more intelligent?
- Could an AI system predict that group average for each face?
For this fixed set, the answer to the second question was yes. Faces ranked relatively high by participants also tended to be ranked relatively high by the AI, and the association was stronger than I expected.
That result does not license the statement “AI can read IQ from faces,” or even the weaker statement “humans can read IQ from faces.” There was no measured IQ and no real person whose intelligence could be inferred. A more accurate description is:
The AI and the participant group responded in partly similar ways to visible features in a controlled set of synthetic portraits.
This is interesting because it may reveal something about shared social perception, cultural learning, visual stereotypes, or the way modern multimodal models reproduce aggregate human judgments. It is also potentially uncomfortable for exactly the same reason. A system can learn to reproduce a stereotype without the stereotype being true.
1. Motivation and scope
People routinely form rapid impressions from faces, even when those impressions are poorly justified. Historical attempts to turn appearance into a science of character or intelligence produced physiognomy, phrenology, racial classification, eugenic measurement, and other harmful pseudoscientific traditions. This history makes precision in the present experiment more than a matter of academic style.
The present study neither rehabilitates those traditions nor tests their central claim. A fictional face cannot validate a theory about intelligence. The experiment instead concerns meta-prediction: can an AI predict how a group of people will answer an explicitly subjective appearance question?
That distinction separates two estimands:
- Not studied: whether appearance predicts a person’s actual intelligence, IQ, character, competence, or worth.
- Studied: whether an AI’s predicted mean appearance rating covaries with the observed mean appearance rating assigned by a particular participant sample to a particular set of synthetic portraits.
The wording matters. Throughout this report, “human score” means the mean perceived-intelligence rating for a face, and “AI score” means the AI’s prediction of that mean. Neither is an intelligence score.
2. Research questions
Primary question
Across the frozen set of 72 synthetic faces, is the AI face-level score monotonically associated with the mean human perceived-intelligence rating?
Secondary descriptive questions
- How closely are the scores related on a linear scale?
- Are AI predictions calibrated to the level and spread of the human means, or do they merely preserve their ordering?
- How stable is the AI when the same face is evaluated in five separate calls?
- How reliable is the aggregate human ordering when the participant ratings are split into independent halves?
Questions the study was not designed to answer
- Can intelligence or IQ be inferred from a face?
- Which facial features cause a person to look intelligent?
- Are participants’ judgments accurate, fair, rational, or desirable?
- Do the observed rankings generalize to natural photographs of real people?
- Would another participant population, model, prompt, image generator, or visual style produce the same result?
- Can the AI predict what a particular individual will answer?
3. Transparency and timing
The project began as a small exploratory video experiment, not as an academic clinical or psychological study. Nonetheless, several safeguards were adopted to reduce opportunities for result-dependent decision-making.
Before the substantive human data were inspected:
- the conceptual estimand was written down;
- the stimulus set was frozen;
- the participant assignment schedule and stopping rule were frozen;
- the inclusion rules and face-level aggregation rule were frozen;
- the primary statistic, permutation test, and participant-cluster bootstrap were specified;
- the analysis code was frozen and checksummed;
- and a separate addendum froze the AI prompt, model configuration, isolation procedure, number of calls, and averaging rule.
These records function as a timestamped, preregistration-like transparency trail. They should not be described as a formal academic preregistration. The work was not peer reviewed, institutionally sponsored, or reviewed by an institutional ethics board. The study design was also developed iteratively through earlier pilots; freezing the main study protects the confirmatory analysis from post-result changes but does not make the entire research program independent of pilot-informed choices.
After the primary result was calculated, individual faces were not immediately browsed for explanations. A small, rule-based set of faces was first selected using numerical categories, the selection was sealed, and qualitative notes were written before the category labels and values were revealed. Those notes are an exploratory audit, not part of the primary evidence.
4. Methods
4.1 Stimuli
The final stimulus set contained 72 AI-generated adult portraits. All images used the same broad photographic shell:
- square, head-and-shoulders framing;
- direct gaze;
- neutral, closed-mouth expression;
- eye-level viewpoint;
- soft studio-like lighting;
- plain gray background;
- and a charcoal crew-neck top.
The prompts excluded text, logos, glasses, hats, jewelry, contextual props, conspicuous makeup, glamour styling, and explicit symbols of education, occupation, intelligence, or status.
Within that shell, the identity descriptions deliberately allowed visible variation in age appearance, presentation, ordered skin-tone target, facial adiposity, face shape, hairstyle, skin texture, and asymmetry. The construction grid contained two presentation targets, three broad age-appearance bands, six ordered skin-tone targets, and two identities per cell, giving 72 faces. These are generation targets, not measured or verified demographic attributes.
The first technically valid generated image for each identity specification was accepted. All 72 were accepted on the first attempt. Faces were not selected or rejected because they appeared attractive, intelligent, unintelligent, likely to produce agreement, or likely to strengthen the hypothesis.
The set was purposively heterogeneous, not randomly sampled from a defined population of synthetic faces. That choice helps prevent a sequence of nearly interchangeable “average” AI portraits, but it also matters for interpretation: a deliberately varied set can create more stable between-face rankings than a narrow or naturally occurring stimulus distribution.
4.2 Participants and recruitment
One hundred and twenty completed participants were recruited through Prolific. The stopping rule was completion of all 120 predetermined assignment blocks, not the attainment of statistical significance.
The analysis included sessions that were consented, completed, and structurally valid with all 19 required ratings. It did not exclude people for unusual answers, fast or slow completion, disagreement with their hidden repeat, weak apparent attention, low variance, or disagreement with the AI. Eleven released or otherwise incomplete sessions existed outside the completed sample; their 85 saved ratings were excluded under the frozen rule and retained for audit rather than silently discarded.
The completed sample contributed 2,280 rating records. Of these, 2,160 were the primary non-repeat ratings and 120 were concealed repeat trials.
Recruitment encountered an operational mismatch between Prolific’s automatic recycling of returned or timed-out places and the website’s deliberately manual release of reserved blocks. The original Prolific export contained many attempts that never created a study session, while some people completed the external study and then returned their Prolific submission. A participant-level reconciliation was performed without consulting face ratings: incomplete reservations were released only after their platform status was confirmed, valid website completions were retained, and a separate seven-place replacement study filled the final seven blocks. Nineteen valid completions were associated with returned Prolific submissions and were designated for full payment by bonus. This incident affected recruitment administration, not the frozen stimuli, block contents, rating question, completion criteria, or stopping rule. The complete reconciliation is retained privately.
4.3 Assignment and presentation
A fixed, phenotype-blind assignment schedule was generated using opaque face identifiers. The scheduling algorithm did not use age, skin tone, adiposity, presentation, or any other visible phenotype when deciding which participants would see which images. This avoided encoding an untested theory of which attributes ought to be balanced.
Each participant rated:
- 18 unique faces; and
- one concealed repeat of one of those faces.
The 19 trials were shown in randomized order, with the repeat separated from its original by at least five trials. Across the completed sample, every face received exactly 30 independent non-repeat ratings. Pairwise face co-occurrence was tightly controlled: any two faces appeared together in between 5 and 10 participant blocks.
Because participants saw only subsets of the stimulus set, the same 30 people did not rate every face. The balanced schedule reduces gross coverage differences, but face means still come from overlapping, non-identical participant groups.
4.4 Human task
Participants were told that the faces were synthetic and fictional. For every image they answered:
Based only on immediate appearance, how intelligent does this fictional person look to you?
Responses used a seven-point verbally anchored scale. The wording intentionally asked about appearance rather than actual intelligence. Participants were debriefed that agreement among raters—or between people and an AI—would not demonstrate accuracy.
The primary human score for each face was the arithmetic mean of its 30 non-repeat ratings. Concealed-repeat ratings were excluded from that score by design and retained for quality and consistency analyses.
4.5 AI instrument
The primary AI instrument was a specific configured system rather than an abstract category called “AI”:
- Codex CLI version 0.146.0-alpha.3.1;
- model gpt-5.6-sol;
- reasoning effort set to xhigh;
- the frozen prompt and structured-output schema;
- and the CLI’s then-current image-processing defaults.
The model was explicitly told that:
- the depicted person was synthetic and had no real intelligence;
- it was not being asked to infer actual IQ;
- its task was to predict the arithmetic mean response from the human study;
- the exact human question and seven-point scale were the target;
- and it should output only the requested structured rating.
Each face was evaluated in a fresh isolated process and context. The file was presented under the generic name “face.webp”; the face identifier was managed outside the model context. No previous faces, human results, prior AI ratings, or study-level comparisons were provided. Browser access, apps, memory, and unrelated tools were disabled.
Five separate calls were made for each of the 72 faces, for a total of 360 calls. Execution order was independently shuffled across runs. All 360 calls succeeded on the first attempt. The primary AI score for a face was the untrimmed arithmetic mean of its five outputs. No call was selected, rejected, calibrated, or replaced according to its value.
The interface did not expose a seed, temperature, or top-p setting, and exact image-detail mode was not explicitly set. Repetition therefore measures the practical stability of this configured instrument, not reproducibility under fully controlled sampling parameters.
4.6 Primary analysis
The unit of analysis was the face, giving 72 paired observations:
- the mean of 30 non-repeat human ratings; and
- the mean of five isolated AI predictions.
The prespecified primary effect was Spearman’s rank correlation, rho. It asks whether faces that receive higher human means also tend to receive higher AI scores without requiring a perfectly linear relationship or identical scale use.
The prespecified null test randomly permuted the mapping between the 72 AI and human face scores 100,000 times. A two-sided Monte Carlo p-value was calculated with the plus-one correction:
(number of permuted absolute correlations at least as large as the observed value + 1) / (100,000 + 1)
Uncertainty due to participant sampling was estimated with 10,000 participant-cluster bootstrap replicates. Entire participants were resampled, preserving each person’s set of ratings and the dependence created by the incomplete-block design. Human face means were recalculated in every replicate and correlated with the fixed AI scores.
This bootstrap is conditional on the 72 chosen faces. It does not treat the faces as a random sample from a broader population of possible images.
4.7 Secondary analyses
The frozen plan treated Pearson correlation, calibration summaries, human reliability, AI run-to-run stability, and visual or feature-level follow-up as secondary or exploratory. These analyses describe the result but do not replace the primary test.
5. Results
5.1 Primary result
Across the 72 faces:
| Quantity | Result |
|---|---|
| Spearman rank correlation | 0.673 |
| Permutations | 100,000 |
| Permutations as or more extreme | 0 |
| Plus-one Monte Carlo p-value | 0.000010 |
| Participant-cluster bootstrap 95% interval | 0.495 to 0.706 |
For this fixed set, the AI’s predicted ordering was strongly associated with the ordering of the participant means. The permutation result says that an association this large would be very unusual under random relabeling of the 72 paired face scores.
It does not say that the hypothesis has a 99.999% probability of being true, that the effect will generalize with that probability, or that any judgment is accurate. The reported p-value is also bounded by the number of permutations: because zero of 100,000 random permutations was as extreme, the plus-one estimate is exactly 1/100,001 rather than a claim of p = 0.
5.2 Scale, spread, and calibration
| Face-level summary | Human means | AI means |
|---|---|---|
| Mean | 4.689 | 4.435 |
| Standard deviation | 0.484 | 0.433 |
| Minimum | 3.667 | 3.600 |
| Maximum | 5.700 | 5.300 |
Pearson’s correlation was 0.665, close to the rank result. A simple linear regression of AI means on human means had an estimated slope of 0.595 and intercept of 1.644.
The central finding is therefore better described as agreement in relative ordering than interchangeability of scores. The AI tended to predict lower means and a narrower range. A high correlation can coexist with systematic level differences and meaningful errors for individual faces.
5.3 Stability of the AI instrument
Across the ten possible pairs of the five complete AI runs, face-level Spearman correlations ranged from 0.920 to 0.962, with a median of 0.933. The median correlation between a single AI run and the human means was 0.664, with a range from 0.638 to 0.673.
Within-face AI variability was usually modest but not absent. The median standard deviation across the five calls was 0.089; the 90th percentile was 0.207 and the maximum was 0.277.
These results make it unlikely that the primary association was created by one unusually fortunate model run. They do not establish stability across model versions, prompting strategies, interfaces, or future deployments.
5.4 Reliability of the aggregate human ranking
When the 30 non-repeat human ratings for every face were repeatedly divided into two groups of 15, the median raw split-half Spearman correlation between the two face rankings was 0.630. The central 95% of the split-half values ran from 0.518 to 0.731. After the conventional Spearman–Brown correction for a full 30-rating aggregate, the median estimate was 0.773, with a central 95% range from 0.682 to 0.845.
This shows that the participant group produced a reasonably stable aggregate ranking, while also showing substantial rating noise. The split-half distribution is a reliability diagnostic, not a confidence interval for the primary AI–human correlation.
5.5 Exploratory controlled reveal
After the numerical analyses, 24 faces were selected under frozen rules into six four-face categories: high agreement/high score, high agreement/low score, AI higher than humans, humans higher than AI, high AI instability, and random controls. The selections were sealed before any face was viewed. Qualitative notes were then recorded without access to the category labels or scores and frozen before unblinding.
This procedure was useful as a disciplined way to generate hypotheses about possible visual cues and failure modes. It was not an independent experiment: the viewer was an informed researcher, the panel was deliberately selected, and the notes were qualitative. No apparent feature in that exercise should be presented as a discovered cause.
6. What the result does show
The narrow empirical claim is strong:
Conditional on this fixed study design, these 72 synthetic images, this Prolific sample, this question, and this configured AI instrument, the AI reproduced a substantial part of the participant group’s face ranking.
Several mundane explanations of the result are disfavored by the records:
- It was not created by stopping data collection when significance appeared; the stopping rule was all 120 blocks.
- It was not created by excluding participants whose answers weakened the result; exclusions were structural and frozen.
- It was not created by choosing among several reported AI runs; all five calls per face were averaged under a frozen rule.
- It was not created by showing the AI the human results or the other faces; calls were isolated and completed before unblinding.
- It was not obviously dependent on one stochastic model run; the five complete run rankings were highly consistent.
- It was not produced by unequal face coverage; every face received 30 primary human ratings.
That is enough to make the association worth taking seriously as a result about shared perception. It is not enough to support a claim about intelligence.
7. Why a correlation of 0.673 may be less sweeping than it first appears
The strongest defense against sensationalism is not to minimize the number. It is to identify exactly what could make the number large.
7.1 It is an agreement signal, not an intelligence signal
The images have no intelligence ground truth. The AI could agree perfectly with every participant and the experiment would still contain zero evidence that the judgments correspond to actual cognitive ability. Shared error, shared bias, and shared stereotype all produce agreement.
This is not a caveat attached to the side of the result. It defines the result.
7.2 Both sides of the correlation are averages
The human value is the mean of 30 judgments; the AI value is the mean of five calls. Averaging suppresses idiosyncratic noise on both sides. The primary correlation therefore compares two stabilized group-level quantities, not one AI answer with one person’s answer.
Aggregate-level correlations can be much larger than individual-level agreement. If individuals share a modest common tendency but disagree noisily, their group mean can still be predictable. A reader should not convert rho = 0.673 into “the AI agrees 67% with people,” and the study provides no estimate of how well the AI predicts a randomly chosen participant.
7.3 The faces were deliberately varied
The stimulus set was designed to avoid 72 nearly identical, generator-default portraits. That was methodologically useful, but it widened visible between-face variation. Correlation depends on the distribution of the cases being correlated: when stimuli span a broad range of cues that both systems use, stable ranking is easier than in a narrow set of similar faces.
The effect might be smaller in a random generator-native sample, a tightly homogeneous sample, candid photographs, or a real-world stream of images. This study did not estimate that.
7.4 The inference is conditional on fixed stimuli
The participant-cluster bootstrap asks how the result changes when participants are resampled. It does not resample a population of faces. The interval of 0.495 to 0.706 therefore should not be read as capturing all uncertainty about a new collection of images.
To make a broad claim about faces, one would need stimulus-sampling uncertainty—ideally using independently generated sets, multiple generators and styles, and a hierarchical analysis treating both participants and images as sampled units. With only this set, a stimulus-specific effect can look statistically precise.
7.5 One generator and one portrait style may create a common visual grammar
All faces came from one generation pipeline and shared a controlled studio format. Synthetic-image systems can encode prompt attributes through recurring textures, shapes, lighting responses, or correlated visual artifacts. Humans and the AI may both be reacting to that common grammar.
The exact images were newly generated, so simple memorization of these files is implausible. But novelty of the files does not rule out familiarity with the generator’s style or with the visual conventions expressed in them.
7.6 The model was asked to imitate the group
The prompt did not ask, “How intelligent is this person?” It gave the AI the human question and scale and asked for the expected arithmetic mean response. This makes the task a direct prediction of collective judgment.
The model has been trained on large quantities of human-produced language and images and may have learned cultural associations between appearance and words such as “intelligent.” Reproducing those associations is a plausible capability of a multimodal model. It requires neither mind reading nor a valid mapping from faces to intelligence.
7.7 Shared stereotypes can create genuine, reproducible agreement
If participants associate particular facial configurations, grooming cues, apparent age, expression details, or photographic features with intelligence, and the model has learned similar associations, a strong correlation follows. The association can be statistically robust and socially meaningful while remaining factually ungrounded.
Indeed, agreement can be evidence of a bias worth studying. It should not be rhetorically promoted into evidence for the stereotype.
7.8 The human target is culturally and procedurally specific
The human score is an average from one recruited online sample answering in one language and interface. It is not “what humans think” in an unrestricted sense. Other countries, age groups, languages, social contexts, or levels of familiarity with synthetic images may rank the faces differently.
Participants were also told that the faces were fictional. That knowledge may change how freely or literally they use appearance-based judgments.
7.9 Human and AI viewing conditions were not identical
Humans rated a sequence of 19 images. Even with randomized order, earlier faces can influence scale use through anchoring, contrast, adaptation, or informal calibration. The AI saw each face alone, with no knowledge of the set.
This asymmetry could weaken, strengthen, or reshape the association. It means the AI predicted an aggregate partly created by contexts it was not shown. A future study could preregister both an isolated AI condition and a context-matched sequence condition.
7.10 Rank agreement is not close numerical prediction
Spearman correlation rewards preserved ordering. It is insensitive to a uniform shift and relatively tolerant of compressed scale use. Here, the AI mean was lower than the human mean and its predictions were less dispersed. The regression slope of 0.595 is another sign of compression.
The model can therefore rank the faces similarly while missing the actual mean by a consequential amount for some images. Correlation should be accompanied by calibration plots and absolute-error measures before using the word “accurate.”
7.11 A correlation is not a percentage
Rho = 0.673 does not mean 67.3% accuracy, 67.3% agreement, or a 67.3% chance of guessing correctly. Squaring a Spearman correlation and calling it “the percentage explained” would also be misleading here. It is a measure of monotonic association between two face-level rankings.
7.12 The small face-level sample is only one part of the uncertainty
There were 120 participants, but the primary statistical unit was the face, so the primary N was 72. The bootstrap interval remains fairly broad. More importantly, simply collecting more ratings of the same faces would eventually estimate these 72 human means very precisely without establishing generalization to new faces.
The next major gain in evidence requires new stimuli and new samples, not only more participants per existing image.
7.13 The result belongs to one model instrument
“AI” is too broad a label for the tested system. The result belongs to one model version, CLI version, prompt, output schema, set of defaults, and date. No second model was designated as a co-primary replication. Other systems may be weaker, stronger, differently calibrated, or sensitive to minor prompt changes.
The five-call procedure addresses within-instrument stochasticity. It does not address model multiplicity or version drift.
7.14 The permutation result is narrow
The very small permutation p-value addresses a null in which the AI and human face labels are exchangeable and unrelated within the fixed dataset. It does not test:
- whether the judgments reflect intelligence;
- whether the model will generalize to new images;
- whether a particular facial feature is causal;
- whether the effect is free from cultural bias;
- or whether every implementation and analytical choice is error-free.
Statistical incompatibility with random relabeling is important, but it is not a universal certificate of meaning.
7.15 Development choices were informed by pilots
The main analysis was frozen before its outcomes were inspected, but the project did not emerge without history. Early pilot work influenced the decision to use more visibly varied synthetic faces and helped refine the technical procedure. Those were sensible design improvements, yet they may also have made shared rankings easier to detect.
This does not invalidate the frozen main result. It means the most compelling next test is an exact prospective replication on a newly generated set whose selection and analysis are fully fixed before any human or AI outcomes exist.
7.16 No feature-level causal claim follows
Many visible properties covary within a face. Apparent age, face shape, skin texture, hairstyle, adiposity, expression remnants, symmetry, and generator artifacts may travel together. A post hoc inspection can suggest stories for almost any high or low score.
The present design was phenotype-blind at assignment, not factorial at generation. It cannot identify which cue caused a rating, whether the same cue affected AI and humans, or whether a cue has the same effect across groups. Feature analysis, if attempted, needs independent coding, prespecified hypotheses, adequate multiplicity control, and new confirmatory stimuli.
7.17 The work has not been independently audited or replicated
Checksums, frozen code, complete call logs, balanced coverage, and disclosed exclusions make the analysis inspectable. They do not substitute for peer review, external code audit, or independent replication. A high-standard public post should expose enough aggregate data and code for others to find errors.
8. What can and cannot be concluded
Supported by this study
- The frozen human face means and frozen AI face means have a substantial positive rank association in this 72-image dataset.
- The association is not plausibly explained by random mismatching of the observed face labels under the permutation null.
- The result is reasonably stable to participant resampling within this fixed stimulus set.
- The configured AI instrument produced highly similar rankings across five independent runs.
- Aggregate human rankings were substantially more reliable than individual ratings would be expected to be.
- The AI captured ordering better than exact scale calibration.
Not supported by this study
- Humans can infer a stranger’s IQ from a face.
- AI can infer IQ, intelligence, personality, competence, or character from a face.
- Any face in the set “really is” intelligent or unintelligent.
- The appearance stereotypes shared by the participants and model are accurate.
- The result applies to real people, candid photographs, other cultures, or all synthetic images.
- A model should be used to evaluate people in education, employment, medicine, policing, credit, dating, or any other consequential setting.
- Any protected or demographic group is more or less intelligent.
- A particular visible feature causes the observed ratings.
The last group of claims would not merely be premature. Several would require an entirely different kind of evidence, and some would pose serious ethical problems even as research objectives.
9. Limitations
For clarity, the main limitations are collected here even where they repeat interpretive points above.
- No ground truth: The faces are fictional and cannot validate perceived intelligence against intelligence.
- Fixed, purposive stimulus set: The 72 faces were constructed for visible heterogeneity, not randomly sampled.
- One generation pipeline: Generator-specific artifacts and visual conventions may drive part of the result.
- One controlled portrait style: Standardization improves internal comparison but limits ecological generalization.
- Face-level N of 72: The primary analysis has 72 units despite 120 participants.
- Participant-only bootstrap: The interval captures participant-sampling variation conditional on the images, not stimulus-sampling variation.
- Aggregate target: Means of 30 human judgments and five AI calls can correlate more strongly than individual responses.
- Convenience online sample: A Prolific sample is not a probability sample of humanity.
- Incomplete blocks: Each face was judged by a balanced but non-identical subset of participants.
- Sequential human context: Humans saw 19-image sequences; the AI saw isolated faces.
- Single AI instrument: Generalization across models, versions, prompts, and interfaces was not tested.
- Partially uncontrolled generation parameters: AI rating calls had no exposed seed, temperature, top-p, or explicit image-detail setting.
- Calibration mismatch: Good ordering coexisted with lower and compressed AI scores.
- Pilot-informed development: The main study was frozen, but prior pilots informed stimulus and procedure choices.
- No causal feature isolation: The design cannot attribute effects to specific facial properties.
- No formal peer review or independent audit: The project remains an independently conducted public experiment.
- No formal ethics-board review: Consent and debrief procedures were used, but the work was not institutionally reviewed.
- Potential social-harm risk: Even accurate reporting can be clipped or reframed as support for physiognomy; presentation choices are part of responsible interpretation.
10. Further research
A useful next program would test robustness rather than immediately search for a biological explanation.
10.1 Exact prospective replication
Generate a new 72-face set under a frozen procedure, recruit a new participant sample, and run the same model instrument without revising the analysis after seeing results. This is the cleanest test of whether the present effect survives new stimuli.
10.2 Random or generator-native stimulus sampling
Alongside a deliberately heterogeneous set, sample faces from a clearly defined generation process without balancing or selecting visible attributes. Comparing the two designs would show how much the effect depends on range construction.
10.3 Multiple generators and visual styles
Repeat the study with independent image generators, prompts, photographic styles, backgrounds, expressions, and levels of realism. A cross-generator result would be less vulnerable to a shared synthetic visual grammar.
10.4 Independent participant populations
Run translated and culturally adapted replications in multiple populations. The goal would not be to find a universal “intelligent face,” but to separate broadly shared judgments from locally learned conventions.
10.5 Multiple AI instruments
Freeze several models and prompts in advance, including a minimal prompt and the present mean-prediction prompt. Treat model as a sampled or crossed factor rather than selecting the best performer after the fact.
10.6 Individual-level prediction
Estimate how well the AI predicts an individual response, not only the group mean. A hierarchical model could separate face consensus, participant tendencies, and residual judgment noise. This would prevent an aggregate correlation from being mistaken for person-level agreement.
10.7 Matched and isolated context conditions
Compare the current one-face-per-call AI protocol with an AI condition that receives the same randomized 19-face sequences as participants. Prespecify tests for anchoring, contrast, and scale calibration.
10.8 Reliability and calibration as primary outcomes
Future work should prespecify absolute error, calibration intercept and slope, and perhaps ordinal predictive scoring alongside rank correlation. A model that orders faces correctly but misses the scale should not be described simply as “accurate.”
10.9 Preregistered cue studies
If specific visual hypotheses emerge, test them through controlled counterfactual image pairs in which one feature is varied while other aspects are held as constant as possible. Use independent blind coding and correction for multiple comparisons. Exploratory visual stories from the present faces should be treated only as hypothesis generators.
10.10 Independent replication and adversarial audit
Publish the prompt, schedule logic, analysis code, face-level aggregates, and provenance hashes where licensing and privacy permit. Invite others to reproduce the result and, especially, to test conditions expected to break it.
None of these extensions would transform perceived-intelligence judgments into IQ measurements. A study using real people and cognitive scores would pose a fundamentally different scientific and ethical question; nothing in the present result provides a shortcut to that question or a reason to revive physiognomic claims.
11. Ethics, disclosure, and responsible communication
Human participation
Participants gave consent, were recruited through a paid Prolific study, were told that the faces were synthetic, and received a debrief clarifying that agreement would not establish accuracy. The public release should report compensation and recruitment settings transparently, verify the final payment reconciliation—including the valid completions routed to bonus payment—and protect participant identifiers.
AI assistance
AI was both the object of measurement and part of the project workflow. The evaluated ratings came from the frozen instrument described above. AI tools also assisted with code, quality checks, planning, and drafting, under human direction and review. This should be disclosed rather than presenting the work as unaided.
Communication rule
The study should never be summarized with a headline or thumbnail that states or strongly implies:
- “AI reads IQ from your face”;
- “This is what an intelligent face looks like”;
- “AI proves physiognomy was right”; or
- “Humans and AI can tell who is smart.”
A defensible short description is:
An AI predicted how a participant group would rank synthetic faces on perceived intelligence. That shows shared appearance judgments, not intelligence detection.
Any video, social post, chart, or thumbnail should keep that distinction visible rather than placing it only in a late disclaimer.
Data sharing
A public replication package should, subject to image licensing and participant privacy, include:
- the full methods and AI addendum;
- the participant-facing wording and scale;
- the phenotype-blind assignment procedure;
- the exact AI prompt and structured output schema;
- analysis code and software-environment notes;
- face-level human means, counts, AI runs, and AI means;
- summary diagnostics and plots;
- inclusion and exclusion counts;
- and cryptographic checksums linking the public files to the frozen artifacts.
Raw Prolific identifiers or other potentially identifying participant information should not be released.
12. Conclusion
This experiment found a clear association between one AI system’s predictions and a participant group’s average perceived-intelligence ratings of 72 synthetic faces. The primary Spearman correlation was 0.673, the result was stable under the prespecified participant bootstrap, and five isolated AI runs produced highly consistent face rankings.
That is a real and interesting result. Its defensible interpretation is also narrow. The AI appears able to reproduce a substantial component of a collective appearance-based judgment in a controlled synthetic stimulus set. The study does not show that intelligence leaves a readable facial signature, that the judgments are accurate, or that the effect generalizes to real people.
The most plausible research direction is therefore not “Can we resurrect physiognomy with a better model?” It is: “What shared visual conventions, learned stereotypes, and aggregation effects allow a model to anticipate human first impressions—and under what conditions do those similarities disappear?”
That question is scientifically interesting without granting an old pseudoscience a single inch.
Appendix A: Exact primary numerical summary
- Completed participants: 120
- Unique faces: 72
- Unique primary ratings per face: 30
- Primary non-repeat human ratings: 2,160
- Concealed repeats: 120
- AI calls per face: 5
- Total valid AI calls: 360
- Primary Spearman rho: 0.6725210321480263
- Permutation draws: 100,000
- Extreme draws: 0
- Plus-one Monte Carlo p-value: 0.00000999990000099999
- Participant-cluster bootstrap replicates: 10,000
- Bootstrap 95% interval: 0.4952228762419479 to 0.7055889262159535
- Pearson r: 0.6651580557361579
- Human mean across faces: 4.6888888889
- AI mean across faces: 4.4347222222
- Human face-mean standard deviation: 0.4836934
- AI face-mean standard deviation: 0.4327995
Appendix B: Artifact and integrity record
The project retains frozen local archives for:
- the synthetic stimuli and generation specifications;
- the assignment schedule;
- completed and released human sessions;
- the full 360-call AI batch, including individual outputs and execution order;
- the frozen primary analysis script;
- the primary result;
- secondary diagnostics;
- and the controlled-reveal selection and notes.
The principal checksum manifests recorded at the time of analysis include:
- human archive manifest: afe98804956381e4c13f04acb3ca01010803e61a845f7acfd97b59db96c6036b
- AI batch manifest: 56d73b5ce2e11f34e349d046539e154546f9b3f4c85b57df6057db5ed42e18c1
- primary-result manifest: 3f39c5bcccf5d89c62c54d212f8ec9a1c8128f8cb89cafc1f55571332a2a5e8b
- secondary-diagnostics manifest: 6e66e442b15958c1a1c2bd82d4bd23b7065d607fc739a4216e3e8cfd3fb73fcc
Suggested citation
Alberto Gemma (2026). Can an AI predict which synthetic faces humans think look intelligent? A controlled exploratory study of shared appearance-based judgments. Public research report.
Leave a Reply