Term

Seed

Info

Originally written in French. Translated by AI — the meaning has been preserved, not the prose.

Short definition

A first batch of sources, small and carefully chosen, that gets iterative research going by producing a first state of knowledge.

Full definition

The word circulates in several traditions: in machine learning, the seed is the value that initialises a random number generator; in snowball literature searching, the seed set or start set is the starting set whose references you follow.

In the articles on this blog, the seed is the initial batch of research steered by the state of the corpus. It is defined by its function — bringing out what is missing — and not by its coverage. It is expected to be reliable, varied, close to the heart of the subject and rich enough for a first conceptual structure to emerge.

Usage in the field

Used when framing research carried out in loops — preparing for a conference, a literature review, a brief given to a research agent — at the point of choosing where to start.

Synonyms and variants

"seed corpus", "initial batch", start set.

Not to be confused with

  • Snowballing — a technique that starts from a source and follows its references and citations. It too starts from an initial set, but the bibliography decides what comes next; here, the diagnosis of the corpus does.
  • representative sample — a subset that reflects the proportions of a population. The seed doesn't try to be representative; it tries to make the gaps visible.

Examples

"Start from a small but carefully vetted seed": a few sources chosen to get going research that will end with 82 sources.

Ambiguities / debates

The line between the seed and the first iteration remains blurred: an initial batch enlarged too early turns into conventional collecting.