Dataset Viewer
Auto-converted to Parquet Duplicate
text
stringlengths
101
92.4k
domain
stringclasses
19 values
pipeline
stringclasses
2 values
style
stringclasses
4 values
**Question:** What is the approximate round-trip communication delay between Earth and Mars at an average distance of 225 million kilometers? **Options:** A. 4 minutes B. 12.5 minutes C. 25 minutes D. 50 minutes --- ### **Why This Question Matters** This question tests your understanding of **communicat...
astronomy
option_level
qa
**Question:** What structure's formation is being studied in relation to star clusters in galactic evolution? Options: A. Galactic halo B. Irregular galaxies C. Spiral arms D. Elliptical cores --- ### 1. Importance/Challenge of the Problem Star clusters (like globular and open clusters) act as "fossil...
astronomy
failure_analysis
qa
# Analysis of the Moon's Orbital Period: A Multiple Choice Question Breakdown ## The Question **Question:** How many days does it take the Moon to complete one orbit around Earth? **Options:** A. 27 days B. 365 days C. 24 hours D. 7 days **Correct Answer:** A --- ## Key Concepts and Principles To...
astronomy
option_level
educational_textbook
# ๐Ÿš€ The Mars Communication Delay: A Space-Time Riddle **Can you calculate the round-trip delay between Earth and Mars? Letโ€™s dive inโ€”and learn how to tackle tricky multiple-choice questions!** --- ## **The Question That Tests Your Cosmic Timing Skills** **Q:** *What is the approximate round-trip communication de...
astronomy
option_level
web_article
# **Why Do Binary Black Holes Merge (or Not)? The Surprising Answer Inside** ## **The Cosmic Dilemma: When Do Black Hole Couples Tie the Knot?** Imagine two black holes locked in a cosmic dance, orbiting each other at breakneck speeds. The question is: *Will they collide within the lifetime of the universeโ€”or drif...
astronomy
failure_analysis
web_article
**User:** Hey, Iโ€™ve got this multiple-choice question here about lunar phases, and Iโ€™m a bit stuck. Can you help me work through it? **Assistant:** Of course! Letโ€™s see the question first. **User:** Okay, here it is: *โ€œWhich lunar phase occurs when the Moon is positioned between Earth and the Sun?โ€* The options a...
astronomy
option_level
conversational_dialogue
**User:** Hi, Iโ€™m working on this astronomy problem, and Iโ€™m a bit stuck. The question is: *โ€œHow does the orbital period of an asteroid in the asteroid belt compare to Earth's year?โ€* The options are A) Shorter than 1 Earth year, B) Longer than 1 Earth year, C) Equal to 1 Earth year, D) Varies widely depending on dista...
astronomy
failure_analysis
conversational_dialogue
# Analysis of the James Webb Space Telescope's Primary Scientific Goal ## Problem Statement **Question:** What is the primary scientific goal of the James Webb Space Telescope? A. To observe the early universe B. To study nearby exoplanets C. To replace Hubble D. To map the Milky Way **Answer:** \boxed{B...
astronomy
failure_analysis
educational_textbook
**Q: What essential nutrient must obligate carnivores like cats obtain through meat consumption as they cannot synthesize it?** Options: A. Taurine B. Vitamin C C. Glucose D. Calcium **A:** The correct answer is \boxed{A}. --- ### **Why this question is important/challenging:** This question tests un...
college_biology
option_level
qa
# **Cracking the Cell Cycle Conundrum: CDK4/6 and the G1 Phase** ## *Why This One Multiple Choice Question Could Boost Your Biology Exam Score* Letโ€™s dive into a question thatโ€™s both tricky and foundational for understanding the cell cycle. Ready? **Question:** *Cyclin-dependent kinases (CDK4/6) are critical fo...
college_biology
option_level
web_article
# Breast Cancer Risk Factors: Analysis and Explanation ## Problem Statement **Question:** What factor is most closely associated with an increased risk of developing breast cancer? A. Higher breast tissue density B. Lower body mass index (BMI) C. Decreased estrogen levels D. Reduced sun exposure **Corre...
college_biology
failure_analysis
educational_textbook
**Q:** What is a likely consequence of a keystone frugivore species experiencing population decline in a rainforest? **A:** The correct answer is **A. Reduced seed dispersal**. --- ### 1. **Why This Problem Matters** Keystone species like frugivores (e.g., monkeys, birds, or bats that eat fruit) play disproport...
college_biology
failure_analysis
qa
# Analysis of Multiple Choice Question: Plant Family Classification of Acacia ## Question **Which plant family does the genus *Acacia* belong to, characterized by leguminous pods and often thorny species?** A. Rosaceae B. Fabaceae C. Asteraceae D. Poaceae **Correct Answer:** \boxed{B} --- ## Key Con...
college_biology
option_level
educational_textbook
**User:** Hey, I have this multiple choice question about the endocannabinoid system, and Iโ€™m a bit stuck. Can you help me work through it? **Assistant:** Of course! Let me hear the question and the options, and weโ€™ll tackle it step by step. **User:** Okay, the question is: *What is the primary role of the endoca...
college_biology
option_level
conversational_dialogue
User: Hi, Iโ€™m working on these multiple-choice questions for my biology exam, but Iโ€™m a bit confused about a couple of them. Let me start with the first one: "Excessive consumption of which substance during pregnancy is directly linked to neural tube defects in the fetus?" The options are A. Caffeine, B. Vitamin A, C. ...
college_biology
failure_analysis
conversational_dialogue
# **Why Do Predators Sometimes Make Prey Less Safe? Unpacking the Hidden Dilemma of Risk Enhancement** ## **The Puzzle: When Survival Strategies Backfire** Imagine youโ€™re a small rodent in the wild. A hawk is circling overhead, so you freeze to avoid detection. But then you hear the rustle of a fox nearbyโ€”now you ...
college_biology
failure_analysis
web_article
# **Unlocking Skincare Science: Niacinamide Concentrations and Why They Matter** ## *Why This Multiple Choice Question Could Change Your Skincare Routine* Letโ€™s dive into a question thatโ€™s both *super practical* and a classic example of how small details in multiple-choice questions can make a big difference in re...
college_chemistry
option_level
web_article
**User:** Hey, I have this question about ionization methods in mass spectrometry, and Iโ€™m a bit confused. The question is: *What ionization method operates by photon-induced electron ejection when photon energy exceeds the analyte's ionization potential?* The options are A. Electrospray Ionization (ESI), B. Atmosphe...
college_chemistry
failure_analysis
conversational_dialogue
**Q: What is the recommended adjustment to the evaporation rate if a sleeve adhesive with a current setting of 8 (on a 1-10 scale) is causing excessive drying speed for a thin PET film?** **A:** The correct adjustment is **A. Decrease by 3 units to 5**. Hereโ€™s why: --- ### **1. Why This Problem Matters** Exc...
college_chemistry
failure_analysis
qa
# Solving the Kernel Ridge Regression Puzzle: Why the Answer Isnโ€™t What You Think! ## The Problem Thatโ€™s Bugging You **Question:** Kernel ridge regression (KRR) is used to approximate which component in computational studies of 1D systems? **Options:** A. Exchange Potential B. Coulomb Potential C. Kinetic ...
college_chemistry
failure_analysis
web_article
User: Hi, Iโ€™m stuck on this question about enzyme immobilization. It says: *Which of the following is a common method for immobilizing enzymes onto magnetic carriers? A. Adsorption, B. Covalent bonding, C. Affinity binding, D. All of the above.* The answer is D, but I want to understand why. Can you walk me through it?...
college_chemistry
option_level
conversational_dialogue
# Analysis of the Reaction Between Carbon Dioxide and Olivine (Mgโ‚‚SiOโ‚„) ## Question **What are the primary products formed when carbon dioxide reacts with olivine (Mgโ‚‚SiOโ‚„) during mineral carbonation?** A. Magnesium carbonate and silica B. Calcium carbonate and quartz C. Sodium bicarbonate and silicon dioxid...
college_chemistry
option_level
educational_textbook
**Question:** What is the maximum number of moles of Hโ‚‚O that can be produced from 4 moles of Hโ‚‚ and 3 moles of Oโ‚‚ in the reaction 2Hโ‚‚ + Oโ‚‚ โ†’ 2Hโ‚‚O? A. 4 B. 6 C. 3 D. 8 --- ### Why is this question important/challenging? This question tests your ability to identify the **limiting reactant** in a chemical...
college_chemistry
option_level
qa
# Phosphorus Yield Calculation: Concept, Analysis, and Common Pitfalls ## Problem Statement **Question:** If a river's annual phosphorus load is 150 tons and its drainage area is 30 kmยฒ, what is the phosphorus yield? A. 5 tons/kmยฒ B. 180 tons/kmยฒ C. 120 tons/kmยฒ D. 0.2 tons/kmยฒ **Correct Answer:** \box...
college_chemistry
failure_analysis
educational_textbook
**Q: What ACID property ensures that a transaction is treated as an indivisible unit, either fully completed or completely undone?** **A:** The correct answer is **Atomicity** (\boxed{A}). **Why this is important**: Atomicity is foundational to transaction reliability. It guarantees that partial changes (e.g., a t...
college_computer_science
failure_analysis
qa
# ๐Ÿ›ก๏ธ The Ultimate Guide to Verifying Software Integrity: A GPG Signature Quiz Breakdown ## ๐ŸŽฏ The Question That Could Save Your System Letโ€™s start with the burning question: **"To verify the integrity of a downloaded software package using a GPG signature, which command should be used first?"** Your options ...
college_computer_science
option_level
web_article
# When a Data Breach Hits: Whatโ€™s the Real Fallout? ## The Question That Stumps Many Imagine this: Youโ€™re taking a cybersecurity quiz and come across this question: **โ€œWhich of the following is NOT a direct consequence of a data breach exposing personal information?โ€** A. Phishing campaigns B. Identity thef...
college_computer_science
failure_analysis
web_article
**User:** Hey, I need help with this multiple-choice question about command-line interfaces. Let me read it out: *In a command-line interface, the 'spawn' subcommand requires two parameters and allows one optional parameter. Which of the following represents the number of required parameters? The options are A. 1, B. 2...
college_computer_science
option_level
conversational_dialogue
# Analysis of HTTP Status Code Question ## Question **What HTTP status code indicates a client-side error due to invalid request parameters?** A. 400 Bad Request B. 500 Internal Server Error C. 200 OK D. 301 Moved Permanently **Correct Answer:** A. 400 Bad Request --- ## Key Concepts and Principles ...
college_computer_science
option_level
educational_textbook
**Question:** What algorithm is most suitable for efficiently finding the shortest path in a grid-based game with weighted tiles? A) Breadth-First Search B) Depth-First Search C) A* Search Algorithm D) Dijkstra's Algorithm **Answer:** \boxed{C} --- ### **Why This Question Is Important/Challenging** ...
college_computer_science
option_level
qa
**User:** Hey, I came across this question about the complexity class of optimal PCB routing. It asks which class it belongs to: P, NP-complete, NP-hard, or Co-NP. The options are A to D, and I thought the answer was A (P), but Iโ€™m not sure. Can you help me work through it? **Assistant:** Of course! Letโ€™s start by u...
college_computer_science
failure_analysis
conversational_dialogue
# The Primary Role of a Computer Network: A Comprehensive Analysis --- ## 1. Introduction to the Problem **Question:** *What is the primary role of a computer network?* **Options (Hypothetical):** A. To connect and facilitate communication between devices, enabling data sharing, resource sharing, and coord...
college_computer_science
failure_analysis
educational_textbook
User: Hey, Iโ€™ve been trying to solve this geometric sequence problem, but Iโ€™m stuck. The question is: What is the next term in the geometric sequence 2, 4, 8, 16? The options are A. 18, B. 24, C. 30, D. 32. I thought the answer was 24, but Iโ€™m not sure. Can you help me figure this out? Assistant: Of course! Letโ€™s st...
college_maths
failure_analysis
conversational_dialogue
# Binary to Decimal Conversion: Understanding the Process and Avoiding Common Errors ## Problem Statement **Question:** What is the decimal equivalent of the binary number \(101010_2\)? **Options:** A. 42 B. 43 C. 44 D. 45 **Proposed Solution:** 45 **Correct Answer:** A. 42 --- ## Key Concepts an...
college_maths
failure_analysis
educational_textbook
# Simplifying Boolean Expressions with De Morganโ€™s Theorem: A Step-by-Step Guide ## Why This Matters in the Real World Boolean algebra might sound like abstract math, but itโ€™s the backbone of digital electronics, computer programming, and even everyday devices like smartphones and traffic lights. For example, thin...
college_maths
failure_analysis
web_article
# Analysis of Ground State Energy in an Infinite Square Well When Width is Halved ## Question **Calculate the ground state energy of a particle in an infinite square well if the width is halved, given the original energy was \( E_0 \).** **Options:** A. \( \frac{E_0}{2} \) B. \( 4E_0 \) C. \( \frac{E_0}{4...
college_maths
option_level
educational_textbook
# The Coin Flip Conundrum: Cracking the Probability Puzzle **Question:** If a fair coin is flipped three times, what is the probability of getting exactly two heads? **Options:** A. 1/4 B. 3/8 C. 1/2 D. 1/3 **Answer:** \boxed{B} --- ## Why This Question Matters (Even If Youโ€™re Not a Math Whiz) Pro...
college_maths
option_level
web_article
End of preview. Expand in Data Studio

QVAC Genesis III

This repository contains the QVAC Genesis III dataset associated with the COLM 2026 paper QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training.

QVAC Genesis III is a large-scale synthetic STEM corpus for efficient language-model pre-training. It contains 159,646,553 educational documents and approximately 191 billion training tokens across 19 curriculum-aligned domains, two complementary generation pipelines, and four educational writing styles.

QVAC Genesis III synthetic-data generation pipeline, from seed acquisition and multiple-choice question generation through student answering, LLM parsing, Option-Level Reasoning, and Failure Analysis

๐Ÿ“‹ Dataset Details

Field Value
Curated by Davide Vitabile, Nikhil Ranjan, Akshay Nambiar, Kamal Kumar Gupta, and Amril Nazir
Organization Tether Data, S.A. de C.V. (Tether AI Research)
Shared by QVAC
Language English
License CC BY-NC 4.0
Scale 159,646,553 documents and approximately 191 billion training tokens
Domains 19 STEM domains
Generation methods Failure Analysis and Option-Level Reasoning
Styles educational textbook, web article, question-answer tutoring, and conversational dialogue

โœจ What Genesis III Adds

QVAC Genesis III uses both successful and unsuccessful answers from an edge-scale student model as generation signals:

  • Failure Analysis (failure_analysis) converts incorrect or non-extractable student responses into corrective lessons. The generated document diagnoses the likely mistake, explains the correct approach, and restates the problem so the result is self-contained.
  • Option-Level Reasoning (option_level) converts correctly answered questions into contrastive explanations that justify the correct option and explain why each distractor is wrong.

Both pipelines render their output in four styles to increase structural and pedagogical diversity.

Under controlled training settings, models trained from scratch on Failure Analysis, Option-Level Reasoning, and their combination substantially improve ARC, GPQA Diamond, and MMLU STEM performance over token-matched Cosmopedia-v2 baselines.

QVAC Genesis III Combined compared with Cosmo-1B and the seven-epoch Cosmopedia-v2 baseline on ARC-Easy, ARC-Challenge, GPQA Diamond, MMLU STEM accuracy, and MMLU STEM Valid Answer Rate

๐Ÿงช Research Checkpoints

Models trained in the paper's controlled from-scratch experiments:

Checkpoint Role Hugging Face
๐Ÿ” Failure Analysis FA data ablation qvac-genesis-iii-qwen3-1.7b-fa
๐Ÿง  Option-Level OL data ablation qvac-genesis-iii-qwen3-1.7b-ol
๐Ÿ”— Combined FA+OL Full Genesis III mix qvac-genesis-iii-qwen3-1.7b-combined

๐Ÿ—‚๏ธ Dataset Configurations

The dataset has one train split and 20 configurations:

  • full (default): all 19 domains
  • one configuration for each individual domain

The configurations reference the same underlying Parquet files; selecting a domain does not create a duplicate copy of the data.

Domains

astronomy college_biology college_chemistry
college_computer_science college_maths college_medicine
college_physics conceptual_physics econometrics
electronic_science geography high_school_biology
high_school_chemistry high_school_computer_science high_school_maths
high_school_physics high_school_statistics machine_learning
professional_medicine

Release identifiers preserve two source-directory names: electronic_science corresponds to electrical engineering, and geography corresponds to high school geography.

๐Ÿงฑ Dataset Structure

Each row contains the final self-contained training document and explicit metadata:

Field Type Description
text string Final synthetic educational document used for pre-training.
domain string One of the 19 curriculum-aligned STEM domains.
pipeline string failure_analysis or option_level.
style string educational_textbook, web_article, qa, or conversational_dialogue.

Generation prompts, generator scratch reasoning, duplicated output fields, and the redundant correctness flag are not included in the public schema. We applied the same record filter used for training, keeping documents whose text is present, non-empty, and longer than 100 characters and whose reasoning_output is present and valid. This filtering removed approximately 1.02 million incomplete or placeholder source records.

Illustrative, abridged example:

{
  "text": "**User:** I have a question about geostationary orbit...",
  "domain": "astronomy",
  "pipeline": "option_level",
  "style": "conversational_dialogue"
}

๐Ÿš€ How to Load

The full corpus is very large. Streaming is strongly recommended unless you intentionally want a local copy.

from datasets import load_dataset

# Stream the complete corpus (the default configuration is "full").
dataset = load_dataset(
    "qvac/GenesisIII",
    split="train",
    streaming=True,
)
example = next(iter(dataset))
print(example)

Load one domain:

from datasets import load_dataset

astronomy = load_dataset(
    "qvac/GenesisIII",
    "astronomy",
    split="train",
    streaming=True,
)

Filter by pipeline or style:

failure_analysis = astronomy.filter(
    lambda row: row["pipeline"] == "failure_analysis"
)

textbook = astronomy.filter(
    lambda row: row["style"] == "educational_textbook"
)

๐Ÿ—๏ธ Dataset Creation

1. Seed acquisition and quality filtering

Seed passages were sampled from FineFineWeb STEM categories and mapped to curriculum-aligned target domains. The Ultra-FineWeb classifier retained high-quality passages.

2. Multiple-choice question generation

QwQ-32B generated self-contained multiple-choice questions with four mutually exclusive options and one gold answer. A format validator rejected malformed generations.

3. Student answering and answer extraction

Qwen3-1.7B-Base produced a free-form answer to each question. CompassJudger-2-32B-Instruct acted as a parser: it extracted one final option or returned a non-answer/ambiguity outcome. It was used to recover the student's committed choice, not to solve the question.

4. Dual-route educational generation

Incorrect and non-extractable responses were routed to Failure Analysis. Correct responses were routed to Option-Level Reasoning. QwQ-32B generated self-contained documents in four styles:

  • educational textbook;
  • web article;
  • question-answer tutoring;
  • conversational dialogue.

The complete methodology and prompt templates are documented in the Genesis III publication.

๐Ÿ›ก๏ธ Content Safety and Personal Information

The complete corpus was scanned before release. The scan covered all source text fields rather than a sample:

  • The toxicity pipeline used Detoxify for broad screening and Qwen3Guard-Gen-8B for contextual review of flagged passages.
  • The PII pipeline detected common personal-data categories, including email addresses, phone numbers, network addresses, physical addresses, and financial identifiers. CompassJudger-2-32B-Instruct reviewed ambiguous findings in context.
  • Confirmed personal identifiers were removed by replacing them with safe, non-resolving synthetic placeholders. The cleaned corpus was then scanned again to verify that no confirmed identifiers remained and that its structure was intact.

The personal-information clean-up modified 379,336 of 160,668,773 source records, approximately 0.24%, removing confirmed identifiers through placeholder substitution. Contextual review found zero toxic passages, so no toxicity-based removals were necessary.

These automated and model-assisted scans can produce false positives or false negatives. Redaction may also affect a small amount of non-personal text that resembles an identifier. The scan addresses toxicity and personal information; it does not establish factual correctness, copyright clearance, or suitability for a particular downstream use. Although automated cleaning, filtering, and deduplication were applied, the dataset was not exhaustively audited for personal, sensitive, offensive, or otherwise harmful content. Users should perform additional screening appropriate to their use of the dataset.

๐Ÿงน Deduplication and Benchmark Decontamination

  • MinHash near-deduplication with 60-gram signatures flagged 47,927 candidate pairs and identified 1,729 unique near-duplicate documents, less than 0.002% of the corpus.
  • AllenAI Decon overlap checks identified 17 verified benchmark-contaminated documents across the corpus.
  • All verified matches occurred in the Failure Analysis data; Option-Level Reasoning had no verified matches.
  • GSM8K test leakage was 0%; MMLU test overlap was 0.005%.

These automatic and manually reviewed checks quantify known overlap but cannot prove that the corpus is free of all benchmark material. The retention or removal status of the identified near-duplicates and benchmark matches must be confirmed before publication.

๐ŸŽฏ Uses

โœ… Direct Use

The dataset is intended for research and education purposes and in connection with those uses, for:

  • pre-training and continual pre-training of language models;
  • research on synthetic educational and STEM data;
  • controlled studies of failure-focused and option-level supervision;
  • curriculum construction by domain, pipeline, or educational style;
  • research on token-efficient training of small language models.

๐Ÿšซ Out-of-Scope Use

The dataset should not be treated as:

  • a factual reference or source of ground truth;
  • a replacement for textbooks, instructors, clinicians, engineers, or other domain experts;
  • human-authored educational material;
  • validated training data for medical, legal, safety-critical, or other high-stakes deployment;
  • evidence that a downstream model is safe or reliable.

Models trained on this corpus require independent domain-specific and safety evaluation before deployment.

โš ๏ธ Bias, Risks, and Limitations

  • The documents are synthetic and may contain hallucinations, reasoning errors, contradictory explanations, or incorrect answer keys.
  • Generator, student, parser, and filtering models can transmit their own biases and systematic errors.
  • Coverage is restricted to English-language STEM and is not representative of broader languages, cultures, disciplines, or educational settings.
  • Domain and difficulty distributions are uneven.
  • Educational styles are generated templates, not naturally occurring classroom interactions.
  • Automated PII and toxicity screening can produce both false positives and false negatives.
  • Redaction placeholders can reduce fidelity or collapse distinct values into the same token.
  • Benchmark decontamination is threshold-based and may miss semantic or paraphrased overlap.
  • Strong benchmark results do not imply suitability for high-stakes applications.

Users should inspect and filter the corpus for their specific use case.

โš–๏ธ Licensing

Licensing Information: This QVAC Genesis III dataset is licensed by Tether Data, S.A. de C.V. under the CC-BY-NC 4.0 (https://creativecommons.org/licenses/by-nc/4.0/legalcode.en). As described in the blog post (https://huggingface.co/blog/qvac/genesis-iii/), the data was generated using FineFineWeb (https://huggingface.co/datasets/m-a-p/FineFineWeb) licensed under the Apache 2.0 license, the Ultra-FineWeb-classifier (https://huggingface.co/openbmb/Ultra-FineWeb-classifier) licensed under the Apache 2.0 license, QwQ-32B (https://huggingface.co/Qwen/QwQ-32B) licensed under the Apache 2.0 license and Qwen/Qwen3-1.7B-Base (https://huggingface.co/Qwen/Qwen3-1.7B-Base) licensed under the Apache 2.0 license.

FineFineWeb data was generated from FineWeb data (ODC-BY 1.0) using data that originated in Common Crawl (Creative Commons terms). The FineFineWeb authors used Qwen2-7B licensed under the Apache-2.0 license, Qwen 2-72B-Instruct licensed under the Tongyi Qianwen license, and a BERT model licensed under the Apache 2.0 license.

Users are responsible for reviewing the applicable licenses and terms of these upstream resources. The dataset is provided as-is, without warranties of factual accuracy, fitness for a particular purpose, or non-infringement.

๐Ÿ“ฎ Copyright Complaints

We will take appropriate action in response to notices of copyright infringement. If you believe your work has been used or copied in a way that infringes your intellectual-property rights, email data-apps@tether.io and identify the copyrighted work and the allegedly infringing material.

๐Ÿ“– Citation

@misc{vitabile2026qvacgenesisiii,
  title         = {QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training},
  author        = {Davide Vitabile and Nikhil Ranjan and Akshay Nambiar and Kamal Kumar Gupta and Amril Nazir},
  year          = {2026},
  eprint        = {2609.19513},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  institution   = {Tether Data, S.A. de C.V. d.b.a. Tether AI Research},
  note          = {Accepted at the Conference on Language Modeling (COLM) 2026},
  url           = {https://arxiv.org/abs/2609.19513}
}

โœ‰๏ธ Dataset Card Contact

Questions and feedback can be submitted through the QVAC organization on Hugging Face. Copyright complaints should be sent to data-apps@tether.io.

Downloads last month
33

Collection including qvac/GenesisIII

Paper for qvac/GenesisIII