American Journal of Public Health Archive (1899–1930)
Dataset Documentation

πŸ“Š AI Training Readiness Report

Overall AI Training Readiness Score: 72 / 100

Executive Summary

This corpus represents a substantial and ethically curated collection of early American public health literature from the Progressive Era.Β 

πŸ“ Dataset Dimensions

MetricValue
Total Records45,696
Unique Titles43,408
Unique Authors636
Unique Issues43 (spanning 1899–1930)
Record Types2 (article: 44,984 / metadata: 712)
Avg. Text Length~720 characters
Text Length Range201 – 5,062 characters

βœ… Strengths (What Lifts the Score)

CategoryScoreNotes
Text Completeness10/10Zero null or empty text fields across all 45,696 records
Ethical Documentation10/10Provenance records, bias audit notices, and clear AI training guidance embedded
Content Richness9/10Deep topical coverage of TB (2,777 records), typhoid (1,605), diphtheria (1,272), sanitation (1,038), vaccination (935), and influenza (602)
Provenance & Licensing9/10Pre-1930 public domain, copyright-cleared, full chain-of-custody documentation
Volume8/1044,984 article-type records is substantial for domain-specific fine-tuning

⚠️ Weaknesses (What Lowers the Score)

CategoryScoreNotes
Author Attribution2/1098.57% of articlesΒ are labeled “Unknown” β€” only 642 records have any author information
OCR Residual Artifacts7/10~4.33% of articles contain detectable artifacts (e.g., “them. to.” for “M.D.”, “JouHn” for “John”, “wlth” for “with”). True artifact rate is likely higher with broader pattern testing
Structural Organization6/10Records appear to be paragraph-level chunks (median 555 chars) rather than reassembled full articles; no explicit publication year column extracted from issue identifiers
Author Field Corruption3/10Where authors ARE present, many show severe OCR distortion (e.g., “B. L. ARMS, them. to., Boston” or “examination of the pancreas, the pylorus…” as an author name)

πŸ“ˆ Text Length Distribution (Article Records)

Length CategoryRecords% of Total
Short (100–300 chars)8,81219.6%
Medium (301–1,000 chars)26,60459.1%
Long (1,001–3,000 chars)9,30420.7%
Very Long (3,000+ chars)2640.6%

πŸ› οΈ Recommendations for Improvement

1. Author Field Remediation (High Impact, +8–10 pts)

  • Cross-reference issue metadata and known AJPH contributor lists from bibliographic databases (e.g., PubMed historical archives, HathiTrust) to backfill authorship.
  • Apply NER (Named Entity Recognition) to the body text to extract author names where they appear in article headers or footers that were OCR’d into the body.

2. Second-Pass OCR Artifact Cleanup (Medium Impact, +3–5 pts)

  • Target the systematic “them. to.” β†’ “M.D.” substitution pattern across the corpus.
  • Apply contextual spell-checking for common OCR confusion pairs (e.g., “wlth”β†’”with”, “Inthe”β†’”In the”, “JouHn”β†’”John”).
  • Quarantine or flag the records where non-text content was mis-parsed into the author field.

3. Structural Enrichment (Medium Impact, +4–6 pts)

  • Extract explicit year/volume columnsΒ from the issue identifier string (e.g., parse “1927-sim_american-journal-of-public-health_19” into year=1927).
  • Reassemble paragraph chunks into full articlesΒ by grouping sequential records with the same title and issueβ€”this would dramatically improve usefulness for long-context LLM training.
  • Add topic/subject tagging using keyword-based or LLM-assisted classification (e.g., “Tuberculosis,” “Sanitation,” “Child Welfare”).

4. Add a Dedicated "Publication Year" Dimension (Low Effort, +2 pts)

  • Enables temporal analysis, time-series training, and period-specific fine-tuning without string parsing at query time.

🎯 Best Use Cases for This Dataset

  1. Domain-Specific Fine-TuningΒ β€” Train or fine-tune LLMs for historical public health, epidemiology, and medical history Q&A
  2. RAG (Retrieval-Augmented Generation)Β β€” Excellent for building a searchable knowledge base of early 20th-century public health policy
  3. Historical NLP ResearchΒ β€” Study the evolution of medical/scientific language from 1899–1930
  4. Comparative Epidemiological AnalysisΒ β€” Train models to compare historical vs. modern pandemic response strategies

Questions?

Please return to the Snowflake Marketplace listing for this dataset to contact us with any questions.

Report prepared June 7, 2026 | Dataset: American Journal of Public Health 1899–1930 | Curated by Devin Media Corp

Scroll to Top