This Soil Microbiome SFA Viral Production Amplicon readme.txt file was generated on 2025-10-22 by Amy Zimmerman GENERAL INFORMATION 1. Title of Dataset: Viral Production 16S rRNA Amplicon Data (raw and processed) 2. Principal Researcher: Name: Kirsten Hofmockel Institution: Pacific Northwest National Laboratory Email: kirsten.hofmockel@pnnl.gov ORCID: 0000-0003-1586-2167 3. Additional Author Contact Information Name: Amy Zimmerman Institution: Pacific Northwest National Laboratory Email: amy.zimmerman@pnnl.gov ORCID: 0000-0001-6709-8274 4. Information about funding sources supporting the data: This program is supported by the U. S. Department of Energy, Office of Science, Office of Biological and Environmental Research, through the Genomic Science Program, under FWP 70880. 5. Geographic location of data collection: 46°15'04"N, 119°43'43"W, Prosser, WA, USA 6. Date of data collection: 2024-06-05 (initial soil collection); 2024-06-11 to 2024-06-13 (collection of incubated soil) DATA & FILE OVERVIEW 1. File List: SoilMicrobiomeSFA_viral-production_amplicon-ASV-counts.csv : Sample specific read counts for each ASV from 16S rRNA amplicon sequencing. SoilMicrobiomeSFA_viral-production_amplicon-metadata.cvs : Sample metadata and total reads from 16S rRNA amplicon sequencing. SoilMicrobiomeSFA_viral-production_amplicon-taxonomy.cvs : Taxonomic assignments for each ASV from 16S rRNA amplicon sequencing. SoilMicrobiomeSFA_viral-production_amplicon-raw/ (DIRECTORY): 120 FASTQ sequencing files, 4 files for each of the 30 samples. File names containing R1/R2 are the forward/reverse reads from the Illumina MiSeq instrument and contain the taxonomic sequence for the samples. Files with I1/I2 are the indexed read files used to demultiplex the dataset on the Illumina MiSeq. 2. Relationship between files, if important: The three CSV files represent different facets of the processed amplicon sequence data. The variables “seqID” and “ASVid” are keys to relate data in these files to each other. All other files are raw sequence files that correspond to the “seqID” in the metadata.csv file. All raw sequence files were processed as a single batch from one Illumina MiSeq flow cell. 3. Additional related data collected that was not included in the current data package: Associated CO2 respiration data, 16S rRNA qPCR data, and bacterial/viral microscopy-based count data from the sample soil incubation experiment are included in a separate zip file. 4. Are there multiple versions of the dataset? No DATA-SPECIFIC INFORMATION FOR: SoilMicrobiomeSFA_viral-production_amplicon-ASV-counts.csv 1. Number of variables: 3 2. Number of cases/rows: 3337291, including header row 3. Variable List: seqID: unique sequencing sample identifier; corresponds to unique sample_cat ASVid: unique identifier for each ASV in format “ASV######” ASV_read_count: sample specific number of sequence reads for specified ASVid 4. Missing data codes: None 5. Specialized formats or other abbreviations used: None DATA-SPECIFIC INFORMATION FOR: SoilMicrobiomeSFA_viral-production_amplicon-metadata.cvs 1. Number of variables: 6 2. Number of cases/rows: 31, including header row 3. Variable List: sample_cat: unique Soil Microbiome SFA sample catalog number in format “SM####” treatment: viral treatment as “field_abund” or “reduced_abund” to indicate whether viruses were added at measured field abundance (1.6 x109 viruses gdw-1 soil) or at 10% of measured field abundance (1.6 x 108 viruses gdw-1 soil) time_h: numeric hours of incubation rep: numeric condition replicate number (none resampled over time), 1-5 seqID: unique sequencing sample identifier; corresponds to unique sample_cat Reads: total number of sequence reads from sample 4. Missing data codes: None 5. Specialized formats or other abbreviations used: gdw == grams dry weight soil DATA-SPECIFIC INFORMATION FOR: SoilMicrobiomeSFA_viral-production_amplicon-taxonomy.cvs 1. Number of variables: 8 2. Number of cases/rows: 111244, including header row 3. Variable List: ASVid: Unique identifier for each ASV in format “ASV######” Domain: Domain level taxonomic assignment Phylum: Phylum level taxonomic assignment Class: Class level taxonomic assignment Order: Order level taxonomic assignment Family: Family level taxonomic assignment Genus: Genus level taxonomic assignment Species: Species level taxonomic assignment 4. Missing data codes: NA (unassigned rank - the ASV can't be confidently assigned to a specific taxon) 5. Specialized formats or other abbreviations used: None METHODOLOGICAL INFORMATION 1. Description of methods used for collection/generation of data: This study aimed to quantify rates of viral production in soil from different viral abundance treatments under conditions as close to natural field soil as possible given the perturbations necessary to manipulate viral abundances. Viruses were removed from soil, then added back to virus-depleted soil to control the initial viral abundances at either 100% (field_abund) or 10% (reduced_abund) of measured field abundance to create treatments with field-relevant or reduced viral infection pressure. Replicates (n=5) of batch incubation jars were harvested every 8 hours for 48 hours to enumerate bacteria and viruses by microscopy and profile bacterial community composition by 16S rRNA amplicon sequencing. DNA was extracted from a subset of frozen soil samples for 16S rRNA amplicon sequencing and qPCR analysis. Because all jars were independently sampled (i.e., no jars were re-sampled through time except for respiration measurements), 3 of the 5 jar replicates with the highest viral abundance were chosen from each of 5 time points (0, 16, 24, 32, and 48 hours) for DNA extraction. For each sample, 250 mg of soil was extracted using the Zymo Quick-DNA Fecal/Soil Microbe Miniprep Kit D6010 per the manufacturer’s instructions. DNA was eluted in 50 µL of Zymo elution buffer. DNA yield was quantified by Qubit DNA High Sensitivity assay (Invitrogen) and quality was checked by NanoDrop (Thermo Fisher Scientific) spectrophotometry. DNA samples were amplified to target the V4 region of the 16S rRNA gene using universal primers 515f (5’-GTGYCAGCMGCCGCGGTAA-3’) and 806r (5’-GGACTACNVGGGTWTCTAAT-3’). A single reaction was performed in 50 uL containing 25 uL Platinum II 2X Master Mix, 2 uL of each primer (0.2 uM final), 19 uL of nuclease free water, and 2 uL of DNA template. The thermocycler conditions consisted of a 3-min hot start at 94°C followed by 27 cycles of 94°C for 45 sec, 50°C for 60 sec, and 72°C for 90 sec before a final extension at 72°C for 10 min. PCR products were cleaned using 1.8X AMPure XP beads and eluted in 40 uL final volume. From the cleaned amplicon product, a second round of PCR to append multiplex barcodes was done. This PCR reaction had a final volume of 25 uL and contained 12.5 uL Platinum II 2X Master Mix, 2.5 uL of each barcoding primer (0.25 uM final), 2.5 uL nuclease free water, and 5 uL of the cleaned amplicon product. The thermocycler conditions consisted of a 3-min hot start at 95°C followed by 8 cycles of 95°C for 30 sec, 55 C for 30 sec, and 72°C for 30 sec before a final extension at 72°C for 5 min. PCR products were again cleaned using 1.8X AMPure XP beads and eluted in 50 uL final volume. Triplicate reactions of Picogreen dsDNA quantification assay were performed using 1 uL of clean amplicon product and measured using a BioTek Synergy H1 plate reader. Based on the Picogreen assay quantifications, 40 ng of DNA per sample was pooled into a 15 mL tube. If a sample did not have enough material to pool 40 ng of DNA, 40 uL of volume was used. Pooled samples were concentrated using the Zymo DNA Clean and Concentrator kit and eluted in a final volume of 100 uL of nuclease free water. The final sample pool was quantified using Qubit High Sensitivity DNA kit and measured at 3.44 ng/uL. Standard library preparation was followed for Illumina sequencing on a MiSeq using a v2 500 cycle flowcell with 10 % PhiX. 2. Methods for processing the data: Demultiplexed data exported from the MiSeq were processed using Qiime2 (version 2021.4). The Dada2 plugin was used to trim, filter, merge, and check for chimeric reads. The forward reads were trimmed at position 13 and truncated at position 200. The reverse reads were trimmed at position 6 and truncated at position 144. Taxonomy at the amplicon sequence variant (ASV) level was assigned using the sklearn method within Qiime2 and Silva database version 138. Contaminant ASV sequences were removed using the decontam R package with a threshold setting of 0.5 using the prevalence model. 3. Instrument- or software-specific information needed to interpret the data: Qiime2: version 2021.4 with the Dada2 plugin and sklearn method within Qiime2 and Silva database: version 138 decontam R package 4. Standards and calibration information, if appropriate: 5. Environmental/experimental conditions: Field Collection Soil was collected from the Irrigated Agriculture Research and Extension Center (IAREC) in Prosser, WA (46.251601, -119.728760) from a managed tall wheatgrass (Thinopyrum ponticum) field under drip irrigation applied at 100% of soil field capacity and described previously. Soil was collected three times to prepare for laboratory-based incubation: soil was collected in April and May 2024 to generate viral inoculum; additional soil was collected in June 2024 to generate the virus-depleted soil/bacterial inoculum and dissolved organic matter solution that were used to establish each incubation jar. Viruses extracted from both April and May samples were combined and concentrated prior to inoculation. On each sampling date, >2 kg soil was collected from 0-15 cm in the edges of 100% irrigation plots and transported back to the laboratory in Richland, WA on ice in a cooler and stored at 4°C overnight before processing. Soil was homogenized through 4 mm sieves after removing rocks and roots. Sieved soil was stored at 4°C until further processing (either preparation of viral inoculum or soil inoculum and DOM solution), which was completed within a week of each sample collection. Soil Incubation Design Soil batch incubations included a field equivalent and a reduced viral abundance treatment, where viruses were added back to the soil/bacterial inoculum at either 100% (1.6 × 109 viruses gdw-1 soil; field_abund) or 10% (1.6 × 108 viruses gdw-1; reduced_abund) of the original viral abundance, respectively. The original viral abundance was determined by epifluorescence microscopy as viruses per gram dry weight (gdw) in collected field soil. All incubations were initiated by adding 65 g air-dried virus-depleted soil/bacterial inoculum to 236 ml wide-mouth Ball mason jars and adding solution to reach a final gravimetric water content of 18%. Soil in each jar was gently mixed with an ethanol-sterilized metal scoopula to distribute moisture evenly. We established 5 replicates of each viral treatment for each sample harvest (0, 8, 16, 24, 32, 40, and 48 hours). Jars were sealed with breathable parafilm to allow gas exchange while excluding potential airborne contaminants and incubated at 25°C. Soil incubations were harvested at seven time points: just after initiation (T0) and after 8, 16, 24, 32, 40, or 48 hours of incubation. At each harvest time, each jar’s soil was mixed with a sterile spatula and subsampled into separate tubes (Olympus Plastics) for viral extraction and enumeration (20g), bacterial extraction and enumeration (10g), and amplicon sequencing (2g). All tubes were immediately flash frozen in liquid nitrogen and stored at -80℃ until further processing. 6. Describe any quality-assurance procedures performed on the data: None 7. People involved with sample collection, processing, analysis and/or submission: Regan McDearis, Sheryl Bell, Sharon Zhao, Evan Warburton, Lupita Renteria, Nicholas Reichart, Amy Zimmerman, Kirsten Hofmockel SHARING/ACCESS INFORMATION 1. Licenses/restrictions placed on the data: This work is marked with CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. The authors do request that you appropriately cite the dataset when referencing or re-using the dataset. 2. Links to publications that cite or use the data: NA 3. Links to other publicly accessible locations of the data: NA 4. Links/relationships to ancillary data sets: None 5. Was data derived from another source? No 6. Recommended citation for this dataset: McDearis, R., A.E. Zimmerman, S.L. Bell, K.S. Hofmockel. 2026. Soil viral production count, respiration, and amplicon data. [Data Set] PNNL DataHub. https://doi.org/10.25584/3653988