The Sorcerer II Global Ocean Sampling ExpeditionNorthwest Atlantic through Eastern Tropical Pacific
Imagine trying to read the ocean without scooping a single microbe into a test tube. That was the audacious idea behind Sorcerer II, a sailboat turned sequencing lab that crisscrossed the Atlantic and into the Pacific. Instead of growing microbes in the lab—where most of them refuse to cooperate—Rusch and colleagues filtered seawater, extracted whatever DNA was there, and sequenced it directly. No petri dishes. Just the ocean, on the page.
The voyage was big and methodical. Forty-one locations, forty-four samples, spaced roughly every couple hundred miles across about eight thousand kilometers. Most of the DNA came from particles between 0.1 and 0.8 micrometers, a size sweet spot rich in free-living bacteria and archaea.
Out of that sweep came a staggering dataset: about 7.7 million reads adding up to roughly 6.3 billion bases. Yooseph and colleagues assembled the reads into about 6.4 million contigs, a nonredundant tapestry of around 5.9 billion bases. It was the largest marine metagenomic collection of its time—more than a fivefold jump over their Sargasso Sea pilot—and it changed the scale of what could be asked.
Scale mattered because the diversity was hiding in the long tail. When the team crunched those reads, most of what they saw had never been cataloged. At a stringent 98 percent identity cutoff, about 85 percent of assembled sequences were unique.
Even among the fragments too diverse to assemble cleanly, a majority were new—around 57 percent. That one-two punch—bigger net, more novelty—meant the ocean’s genetic library didn’t just expand; it shifted our center of gravity from cultivated isolates to environmental reality.
Turning raw fragments into biological meaning took a careful, almost artisanal assembly strategy. The primary build used the Celera Assembler with strict settings, insisting on 98 percent overlap identity to avoid stitching together slightly different strains into Franken-genomes. That caution came at a cost.
Only about 9 percent of reads landed in scaffolds longer than 10 kilobases, and just over half of all reads—53 percent—stayed unassembled as singletons. Yooseph’s team probed those parameters by lowering the identity threshold: assembly lengths grew as they relaxed to 94 percent identity. Below 90 percent, assembly stalled or got worse, a clear sign that real polymorphism—not merely noise—was colliding in the overlaps.
Could longer DNA fragments help? In one standout sample, GS-33, they built a fosmid library with inserts averaging about 36 kilobases. That extra reach changed the picture: the largest scaffold ballooned from roughly 70 kilobases to about 1.25 megabases.
The largest contig jumped to around 427 kilobases when combined with a high-stringency assembly. It was like handing a weaver longer threads—they could span the knotted bits—though the overall tapestry was still complex.
They also tried something bold: an “extreme” assembly that grew contigs out from seeds without relying on mate pairs, always chasing the best overlap at each end. In the most abundant lineages, that approach produced whoppers—contigs up to roughly 900 kilobases. They aligned almost end-to-end with the genomes of Prochlorococcus MIT9312 and Pelagibacter ubique HTCC1062.
Tempting, right? But there was a catch. Those aggressive builds carried more artifacts—chimeras, false consensuses—so the researchers treated them as exploratory, not gospel.
For most analyses they returned to the conservative primary assembly. They also made the whole toolkit—assemblies, reads, and new visualization tools—available through the CAMERA database and public archives so others could inspect the seams.
The real breakthrough in making sense of this diversity was conceptual as much as computational: fragment recruitment. Picture a reference genome along the x-axis and, on the y-axis, the percent identity of every environmental read that aligns to each position. Color the dots by sample.
What you see—especially in Pelagibacter, Prochlorococcus, and Synechococcus—are horizontal bands: dense corridors of reads at, say, 85 to 95 percent identity that tile across most of the genome. Each band is a closely related subtype. Within a band, the reads are similar enough to map coherently but different enough—on average, about 3 to 5 percent nucleotide divergence—to reveal microdiversity.
Those bands tell stories. In Pelagibacter, the highest-identity reads cluster from temperate samples, with lower-identity reads from more distant waters fanning beneath. Across the whole dataset, five genera dominated recruitment—Pelagibacter, Prochlorococcus, Synechococcus, Burkholderia, and Shewanella.
Together, they accounted for roughly half of all recruited reads but only about 15 percent of all reads, a reminder that recruitment highlights the abundant and the knowable while much else hums in the background. From a practical standpoint, Yooseph and colleagues found you don’t need to sequence the ocean to exhaustion to see patterns; as few as ten thousand reads could reliably compare surface-water samples.
Now, zoom in. Those tidy bands don’t mean monolithic genomes. They’re more like braided streams.
Hypervariable islands—genomic neighborhoods rich in integrases and recombinases—punctuate the landscape. Gene content in these islands shifts from clone to clone, often in biome-specific ways, and the most dynamic segments simply refuse to assemble. That’s why, even with seed-based extensions targeted at known ribotypes, you get stretches of 100 kilobases or more that are solidly syntenic.
They are interrupted by gaps where the diversity is too tangled to resolve. Seed-based assembly, by the way, is a clever trick: start from a read mated to a telling marker—say, a SAR11 16S fragment—and walk outward along overlapping reads. In one experiment, they seeded 348 such assemblies and picked 24 that were clearly independent.
They used those consensus segments as fresh references to watch different SAR11 subtypes sort by geography.
Structural variation was another question: are these environmental populations constantly flipping and shuffling their genomes? In Prochlorococcus MIT9312-like populations, big rearrangements were surprisingly rare. The team used mate-pair metadata—tracking “missing mates” that suggested a break in synteny—to search for inversions and translocations larger than about 50 kilobases.
They estimated a rate of roughly one such event per 2.6 megabases of sequence. In other words, the core architecture holds. Most of the action is in small-scale polymorphisms and those hypervariable islands, not wholesale genome rewiring.
Taxonomy offered a second, complementary lens. By extracting sixteen S ribosomal RNA genes from the metagenome, the team built a census of who was there: four thousand one hundred twenty-five full or partial sixteen S sequences clustered at 97 percent identity into 811 ribotypes. About 48 percent of those ribotypes matched entries already in public databases; the rest expanded the tree.
Sixteen percent of ribotypes—and a smaller fraction of actual sequences, about 3.4 percent—were more than 10 percent divergent from anything known, pointing to novelty at family-level depths and beyond.
Those ribotypes trace a geography. Samples separated into temperate and tropical cohorts, with some expected outliers: a hypersaline pond known as GS33, a freshwater site GS20. But the taxonomy wasn’t a strict map overlay.
Some Caribbean samples looked taxonomically similar to eastern Pacific ones, yet their functional gene content diverged. This brings us to the second layer: what these communities can do.
Using TIGRFAM profiles—curated protein families—the team contrasted functional repertoires across water masses. The temperate Atlantic carried a strong phototrophic signature: components of Photosystem II and Photosystem I were markedly enriched, with the 44 kilodalton PSII reaction center subunit and the core PSI proteins PsaA and PsaB standing out. Carbon storage and processing also popped, including glycogen and starch biosynthesis.
Phosphate acquisition cut sharply across the transect. Phosphate-binding and transport proteins were more abundant in the temperate cluster than the tropical one, consistent with nutrient constraints in those coastal waters. Even within the tropics, an Atlantic to Pacific split appeared: Caribbean sites showed more of the PstS gene and PstABC-type phosphate uptake genes than their eastern Pacific counterparts.
That hinted at regional differences in phosphate availability that the microbes were gearing up to exploit.
Proteorhodopsins were the showstopper. These are light-driven proton pumps—molecular solar panels—that can supplement energy budgets in what we usually think of as heterotrophs. Across the Global Ocean Sampling dataset, the team found two thousand six hundred seventy-four putative proteorhodopsin genes, and in one thousand eight hundred seventy-four of them, the amino acid that tunes light absorption could be read directly.
That single residue tracks the color of light a protein prefers: leucine skews green, glutamine skews blue. The pattern snapped into place with geography. The green-tuned variant was enriched along the temperate Atlantic near the U.S.–Canada coast and even in the freshwater site, while the blue-tuned form dominated most open-ocean samples.
The moderate class showed up across many environments. And this wasn’t just a quirk of one lineage. About a quarter of proteorhodopsin reads rode in on SAR11 fragments, and phylogenetic reconstructions suggest those SAR11 variants arose independently at least twice.
As Rusch and colleagues put it, proteorhodopsins blur the old heterotroph–autotroph line; light can feed respiration and growth in far more players than we thought.
Not every functional trait marches with phylogeny. The phosphate-binding gene PstS and its neighboring regions showed a bimorphic pattern that could vary within a single subtype, with limited concordance to the sixteen S tree. Proteorhodopsin tuning, by contrast, often tracked lineage boundaries, yet still showed convergent solutions across distant branches.
That mix—lineage-tethered traits here, ecotype-sliced traits there—is exactly what you’d expect in a vast, patchy ocean where both ancestry and environment steer evolution.
All of this rested on a mountain of proteins. The Global Ocean Sampling catalog now includes about six point one million annotated proteins, including thousands of new families, and the data were posted openly to the CAMERA database and the National Center for Biotechnology Information for anyone to mine. Fragment recruitment became a community tool, not just a figure in a paper.
The team stressed its limits along with its power. Under lenient criteria, around 70 percent of reads could be recruited to references; under stringent ones, closer to 30 percent. Extreme assemblies unlocked long contigs in abundant taxa, but they needed to be treated with suspicion.
Hypervariable islands were—and are—the ghosts in the machine, ducking comprehensive assembly and reminding us that population structure lives at multiple scales.
So what did we learn, big picture? First, metagenomics at ocean scale is not just a cataloging exercise. It rewires questions.
Instead of asking “what species are here?” we can ask “how many subtypes travel together?” and “which functions move with them?” Second, abundant lineages in the sea aren’t neat clones; they’re braided populations with conserved backbones, hypervariable islets, and a handful of ecologically decisive genes. And third, the environment leaves its fingerprints everywhere: in the color of light proteins absorb, in the way phosphate is hauled across membranes, and in the quiet scarcity of large-scale genome rearrangements.
Where does it go from here? The Global Ocean Sampling team pointed to two practical fronts. One is methodological—seed-based assembly and richer long-read or long-insert data to stitch through the knotty bits without inventing chimeras.
The other is comparative—expanding the reference set so fragment recruitment can pull more of the dark matter into view. But the core move has already paid off. By sequencing the sea directly, Rusch, Yooseph, and colleagues turned a blue expanse into a legible, if still challenging, genetic landscape.
We can now watch populations, not just species; functions, not just names. And that, for the ocean and for microbiology, is a sea change.
Related lectures
- A global inventory of small floating plastic debris
- Estimating Global “Blue Carbon” Emissions from Conversion and Degradation of Vegetated Coastal Ecosystems
- Interannual variability in global biomass burning emissions from 1997 to 2004
- Adaptation, Plasticity, and Extinction in a Changing Environment: Towards a Predictive Theory
- Regional Decline of Coral Cover in the Indo-Pacific: Timing, Extent, and Subregional Comparisons
- Effects of Roads on Animal Abundance: an Empirical Review and Synthesis