American Journal of Public Health Archive (1899β1930)
Dataset Documentation
π AI Training Readiness Report
Overall AI Training Readiness Score: 72 / 100
Executive Summary
This corpus represents a substantial and ethically curated collection of early American public health literature from the Progressive Era.Β
π Dataset Dimensions
| Metric | Value |
|---|---|
| Total Records | 45,696 |
| Unique Titles | 43,408 |
| Unique Authors | 636 |
| Unique Issues | 43 (spanning 1899β1930) |
| Record Types | 2 (article: 44,984 / metadata: 712) |
| Avg. Text Length | ~720 characters |
| Text Length Range | 201 β 5,062 characters |
β Strengths (What Lifts the Score)
| Category | Score | Notes |
|---|---|---|
| Text Completeness | 10/10 | Zero null or empty text fields across all 45,696 records |
| Ethical Documentation | 10/10 | Provenance records, bias audit notices, and clear AI training guidance embedded |
| Content Richness | 9/10 | Deep topical coverage of TB (2,777 records), typhoid (1,605), diphtheria (1,272), sanitation (1,038), vaccination (935), and influenza (602) |
| Provenance & Licensing | 9/10 | Pre-1930 public domain, copyright-cleared, full chain-of-custody documentation |
| Volume | 8/10 | 44,984 article-type records is substantial for domain-specific fine-tuning |
β οΈ Weaknesses (What Lowers the Score)
| Category | Score | Notes |
|---|---|---|
| Author Attribution | 2/10 | 98.57% of articlesΒ are labeled “Unknown” β only 642 records have any author information |
| OCR Residual Artifacts | 7/10 | ~4.33% of articles contain detectable artifacts (e.g., “them. to.” for “M.D.”, “JouHn” for “John”, “wlth” for “with”). True artifact rate is likely higher with broader pattern testing |
| Structural Organization | 6/10 | Records appear to be paragraph-level chunks (median 555 chars) rather than reassembled full articles; no explicit publication year column extracted from issue identifiers |
| Author Field Corruption | 3/10 | Where authors ARE present, many show severe OCR distortion (e.g., “B. L. ARMS, them. to., Boston” or “examination of the pancreas, the pylorus…” as an author name) |
π Text Length Distribution (Article Records)
| Length Category | Records | % of Total |
|---|---|---|
| Short (100β300 chars) | 8,812 | 19.6% |
| Medium (301β1,000 chars) | 26,604 | 59.1% |
| Long (1,001β3,000 chars) | 9,304 | 20.7% |
| Very Long (3,000+ chars) | 264 | 0.6% |
π οΈ Recommendations for Improvement
1. Author Field Remediation (High Impact, +8β10 pts)
- Cross-reference issue metadata and known AJPH contributor lists from bibliographic databases (e.g., PubMed historical archives, HathiTrust) to backfill authorship.
- Apply NER (Named Entity Recognition) to the body text to extract author names where they appear in article headers or footers that were OCR’d into the body.
2. Second-Pass OCR Artifact Cleanup (Medium Impact, +3β5 pts)
- Target the systematic “them. to.” β “M.D.” substitution pattern across the corpus.
- Apply contextual spell-checking for common OCR confusion pairs (e.g., “wlth”β”with”, “Inthe”β”In the”, “JouHn”β”John”).
- Quarantine or flag the records where non-text content was mis-parsed into the author field.
3. Structural Enrichment (Medium Impact, +4β6 pts)
- Extract explicit year/volume columnsΒ from the issue identifier string (e.g., parse “1927-sim_american-journal-of-public-health_19” into year=1927).
- Reassemble paragraph chunks into full articlesΒ by grouping sequential records with the same title and issueβthis would dramatically improve usefulness for long-context LLM training.
- Add topic/subject tagging using keyword-based or LLM-assisted classification (e.g., “Tuberculosis,” “Sanitation,” “Child Welfare”).
4. Add a Dedicated "Publication Year" Dimension (Low Effort, +2 pts)
- Enables temporal analysis, time-series training, and period-specific fine-tuning without string parsing at query time.
π― Best Use Cases for This Dataset
- Domain-Specific Fine-TuningΒ β Train or fine-tune LLMs for historical public health, epidemiology, and medical history Q&A
- RAG (Retrieval-Augmented Generation)Β β Excellent for building a searchable knowledge base of early 20th-century public health policy
- Historical NLP ResearchΒ β Study the evolution of medical/scientific language from 1899β1930
- Comparative Epidemiological AnalysisΒ β Train models to compare historical vs. modern pandemic response strategies
Questions?
Please return to the Snowflake Marketplace listing for this dataset to contact us with any questions.
Report prepared June 7, 2026 | Dataset: American Journal of Public Health 1899β1930 | Curated by Devin Media Corp