The AI Picbreeder Experiment: In Search of Automatic Open-Endedness

The AI Picbreeder Experiment

Best Paper (Complex Systems), GECCO 2026.

Sam Earle NYU

Kai Arulkumaran Sakana AI

Andrew Dai Independent

Akarsh Kumar MIT

Julian Togelius NYU

Sebastian Risi Sakana AI

Jul 6, 2026

Paper

Code

Data

* Work done during an internship or residency at Sakana AI.

Can AI agents be creative? It’s a question on the lips and fingertips of many.

Sure, AI is the new medium, but could it ever be the one making the message?

We ask whether the agents we use today as tools might be re-purposed to spontaneously speak among themselves.

The end result would be a model organism of cultural production.

What should we ask of the agents to make this real? The tension is that we cannot ask for anything at all—it has to be left up to them—yet modern agents demand our asking by design. They are trained, evaluated, and orchestrated as goal-following entities.

But creativity elides planning, and the end products of cultural processes are not conceived of at the outset but forged in serendipity. This is true not only of the arts, but of math and science as well, where the discovery of new questions is at least as important as their resolution.

To this end, we envision agents capable of intentional aimless wandering—agents that can surprise themselves, and in response throw their plans and preconceptions out the window in pursuit of the unexpected.

To study whether AI systems might have this capacity, we turn to a minimal substrate where we know it to be possible—even necessary. Picbreeder was a website which gained a modest cult following in the 2010s. Here, you might imagine us trying to re-animate it with something like the mechanical ghosts of its past users: replacing humans with large models trained on their collective output, and asking if these models are capable of the same level of creative discovery as we were.

At picbreeder.org, users participated in a process of generating collaborative art by interactively evolving images. The images began abstract, but over multiple sessions with different users, genealogies complexified and familiar forms emerged, often by surprise. (Image retrieved via the Internet Archive.) (Click to access an homage to the original site, populated with aggregate results over multiple replays of AI-driven reproductions of the collective human event.)

Before frontier models allowed us to speak into existence a vast distribution of possible images, Picbreeder had users breed images by representing them indirectly. Each image in Picbreeder is encoded as a Compositional Pattern-Producing Network (CPPN), an initially small, protean neural network that takes as input some coordinates in pixel space, and outputs that pixel’s color. Because they can take arbitrary continuous coordinates as input, these networks effectively encode infinite-resolution images. Elsewhere, they’ve been used to represent videos, terrain in game-like environments, and even the weights of larger downstream neural networks with more regular topologies.

Compositional Pattern-Producing Networks (CPPNs) At every pixel, x, y coordinates, and radius from center r are fed into the network, which returns hue, saturation and value (HSV), i.e., a color. In this way, each CPPN encodes an image of arbitrary resolution.

In practice, we pass all the input coordinates through the network at once. Each node in the network has some grid-shaped activation we can render as a greyscale image, and at the network’s output, three such outputs are combined to produce a single color image.

Playing with the weights for a moment, you get the sense that building target images by building CPPNs from scratch out of nodes and edges would be a thankless endeavor. Instead, in Picbreeder, images are evolved. This is done interactively with the user in the loop: from a grid of random initial CPPNs (which generally look like smooth blobs or gradients), a user selects one or several which they’d like to grow further.

These individuals (each little brain containing a single picture) are then bred together and mutated to produce a new set of candidate images, and the process repeats. Mutation and crossover operations are drawn from the Neuroevolution of Augmenting Topologies (NEAT) algorithm, which unlike traditional gradient descent not only optimizes the weights of the network, but also its shape or topology in terms of the layout of nodes and edges.

Breeding console. Click to select · space to evolve · z to undo · r to reset · p for new parent.

In this way, images are created not by design but by a prolonged process of natural selection. And as the human in this loop, it’s nearly impossible to plan for a particular long-term goal: for example to imagine an image in advance then successfully breed it. Not within this interaction scheme, not on any reasonable kind of timescale, at least.

But Picbreeder shows us that these kinds of achievements can be attained anyway if the process of interactive evolution is distributed across a large number of users. Here, the artifacts materialize by accident, appearing to the user as objectives-in-retrospect.

In other words, the Picbreeder users played with the system, wandering, and generated art not by conceiving it against marble or canvas, but by catching it like a fish from a stream. In the AI Picbreeder experiment, we ask: can AI catch similar fish? And if it can, then how does it achieve this? Then we can try to find any parts of its looking at and moving in the water that matter to its success.

AI Picbreeder: mimicking a pared-down version of human interaction with the original Picbreeder website, we use VLM instances to simulate users, growing unbounded archives of collaborative art produced by agents under varying conditions. In parallel, VLM agents interactively evolve and evaluate candidates saved to an online archive. As the archive grows unboundedly, candidates are selected according to their past evaluation scores and their placement in the phylogenetic tree (recency, number of offspring).

Evaluation metrics. Visual and Semantic Coverage embed images and (VLM-generated) captions to a shared embedding space and measure each archive’s propensity to cover that space in terms of k-covering radii. Semantic Recall pre-defines a set of nouns corresponding to images we might hope to recover, maps images and labels to a joint text and image embedding space, and measures the extent to which generated images approximate the set of labels, via max per-label cosine similarity.

We also develop a choice set of knobs for the system. We’re inspired by the idea that serendipity requires chance, a kind of local narrative flow, and a broader personal context; and design interventions intended to serve as simplified computational models of these effects. In particular, we inject random noise into the agents’ decision making process, we play with their memories—erasing everything but the present moment or overloading them with a full view of everything they’ve ever seen—and we seed them with simple personality traits intended to indirectly influence their interaction with the system.

Quantitative results. The human archive emerges as a strong upper bound relative to our VLM-driven experiments, where we measure the effects of noise, memory and agents on Picbreeder. The random baseline provides a lower bound. Among VLM experiments, agent diversity via the injection of subtle breeder system prompts is particularly effective in closing the VLM-human gap. A minimal amount of memory is crucial in reducing repetitive actions while avoiding context bloat in this token-hungry multimodal regime. A bit of noise in the selection process can also encourage exploration. Uniformly random branching produces maximally balanced trees in expectation.

Sam Earle, Kai Arulkumaran, Andrew Dai, Akarsh Kumar, Julian Togelius, Sebastian Risi, "In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models", GECCO 2026.

BibTeX citation:

@inproceedings{earle2026picbreedervlm,
  title     = {In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models},
  author    = {Earle, Sam and Arulkumaran, Kai and Dai, Andrew and Kumar, Akarsh and Togelius, Julian and Risi, Sebastian},
  booktitle = {Proceedings of the Genetic and Evolutionary Computation Conference (GECCO '26)},
  year      = {2026}
}