JAMA Surgery Dataset 1911 to 1930
Documentation

Overview

The JAMA Surgery Dataset (1911–1930) is a curated collection of historical surgical literature sourced from the Journal of the American Medical Association (JAMA) Surgery archives. This dataset has been professionally processed and structured to support modern data workflows, including machine learning, natural language processing, and historical analysis.

All content originates from verified public domain sources and has been transformed into a clean, machine-readable format suitable for AI training and research applications.

Dataset Structure

The dataset is organized at the article level, with each row representing a structured record extracted from an individual publication.

Fields

  • file_name
    The original source file name corresponding to the archived publication.
  • author
    Extracted author name(s) associated with the article.
  • issue
    Publication issue and date (e.g., “1921-November”).
  • text
    Cleaned and structured article content, processed for readability and downstream AI use.

Data Processing Pipeline

This dataset has undergone a multi-step transformation process to ensure quality, consistency, and usability:

  1. Source Acquisition
    Public domain archival materials were collected from verified repositories.
  2. OCR Processing
    Raw documents were converted into machine-readable text using high-quality OCR.
  3. Data Cleaning
    • Removal of OCR artifacts and noise
    • Correction of formatting inconsistencies
    • Standardization of structure
  4. Normalization
    Metadata fields (author, issue, file_name) were extracted and aligned into a consistent schema.
  5. Validation
    • Manual spot-checking for accuracy
    • Structural validation for consistency
    • Sampling for quality assurance

Data Quality

  • Professionally cleaned and structured
  • Reduced OCR noise and formatting errors
  • Consistent schema across all records
  • Suitable for production-level AI pipelines

Intended Use Cases

This dataset is designed to support a wide range of applications:

  • Training domain-specific medical and surgical language models
  • Retrieval-augmented generation (RAG) systems
  • Historical surgical research and trend analysis
  • Semantic search and knowledge extraction
  • Academic and educational use

Ethical and Legal Considerations

All materials included in this dataset are:

  • Verified as public domain
  • Free from copyright restrictions
  • Ethically sourced and processed

This dataset is intended to provide a high-integrity alternative to unverified, web-scraped training data.

Access and Integration

The dataset is delivered via Snowflake and can be queried directly using standard SQL.

©2026 Devin Media Corp. All rights reserved.

Scroll to Top