JAMA Pediatrics Clinical Dataset (1911–1930)

Overview

This dataset contains over 28,000 structured records from the Journal of the American Medical Association (JAMA) Pediatrics publications spanning 1911–1930.

The dataset has been professionally curated, cleaned, and structured for use in AI training, natural language processing (NLP), and historical research applications.

Dataset Structure

Each record includes the following fields:

– **file_name**: Original source file identifier
– **author**: Extracted author name (where available)
– **issue**: Publication issue (formatted as YYYY-Month)
– **text**: Full cleaned text content of the article or entry

Data Format

– Source format: JSONL (line-delimited JSON)
– Storage format: Snowflake table
– Encoding: UTF-8
– Queryable via SQL

Coverage

– Publication: JAMA Pediatrics (historical)
– Time range: 1911–1920
– Total files: 120+
– Total records: 28,000+

Processing Methodology

This dataset was processed using a multi-step pipeline:

1. OCR extraction from historical source documents
2. Cleaning of OCR artifacts and formatting inconsistencies
3. Structured parsing into JSONL format
4. Extraction of metadata fields (author, issue)
5. Validation and ingestion into Snowflake

Data Quality Notes

– OCR errors have been minimized but may still exist in rare cases
– Author fields may be missing or marked as “Unknown” where not available
– Text formatting reflects original historical structure where appropriate

Ethical Considerations

This dataset consists of pre-1930 public domain material.

– Content reflects the medical knowledge and societal context of its time
– Historical biases, terminology, and perspectives may be present
– This content is provided for research and AI training purposes only

Users should not interpret historical language as reflective of modern medical standards or ethical norms.

Intended Use Cases

– Large Language Model (LLM) training and fine-tuning
– NLP model development
– Historical medical analysis
– Dataset benchmarking and evaluation

Licensing & Provenance

– Source: Verified pre-1930 public domain publications
– Curated by: Devin Media Corp
– Ethical framework: Foundation for Ethical AI

Each record includes embedded provenance and bias audit notices to support responsible AI development.

Access

This dataset is available as a Snowflake-native table and can be queried directly using SQL.

©2026 Devin Media Corp. All rights reserved.

Scroll to Top