Finance AI: Historical Banking & Credit Archive (pre‑1930)
Dataset Documentation

Overview:

This dataset is a substantial, well-structured historical financial corpus comprising over 2 million records drawn from 10 authoritative financial publications spanning from 1843 to the early 1930s. It has undergone rigorous OCR processing and paragraph reconstruction, resulting in excellent core data completeness.

📊 AI Training Readiness Report:
*Overall AI Training Readiness Score: 84 / 100

Dataset readiness report prepared by Dataset IQ.

Dataset Profile

MetricValue
Total Records2,016,529
Article Records1,992,593
Metadata Records23,936
Distinct Publications10
Distinct Issues7,779
Temporal Span~1843 – 1930s
Avg. Text Length846 characters
Median Text Length571 characters

Source Publications

The dataset aggregates content from 10 high-quality historical financial periodicals:

  • The Economist (UK — global finance & trade)
  • Commercial & Financial Chronicle (US — capital markets)
  • Federal Reserve Bank of New York Monthly Review (US monetary policy)
  • Citibank Monthly Economic Letter (commercial banking commentary)
  • Lloyd’s Bank Review (UK banking perspective)
  • Trusts & Estates (fiduciary & trust administration)
  • Business Credit (commercial lending)
  • Business Month (general business economics)
  • Taxes (fiscal policy & taxation)
  • The Annalist: A Magazine of Finance & Commerce (financial analysis)

Scoring Breakdown

CategoryScoreNotes
Volume & Scale19/202M+ records is exceptional for domain-specific AI training
Data Completeness16/200% null text, 0% null issues; author field is 99% “Unknown” (expected for historical periodicals)
Text Quality & RAG Readiness16/20Pre-chunked paragraphs at ideal retrieval lengths; no empty texts; min length 201 chars ensures meaningful content
Structural Consistency17/20Clean schema with type separation, consistent issue identifiers encoding year + publication + date
Domain Richness16/2010 sources across banking, securities, monetary policy, taxation, trust law; 86+ year temporal breadth

Strengths ✅

  • Zero null/empty text fields — every record contains substantive content
  • Pre-chunked paragraph structure — the median length of 571 characters is ideal for RAG retrieval without further splitting
  • Rich domain vocabulary — 30% of articles reference stocks/bonds, 12.6% discuss banking, 4.8% cover credit markets
  • Temporal depth — spans the formation of the Federal Reserve (1913), WWI-era finance, and the lead-up to the Great Depression
  • Publication diversity — covers US, UK, commercial, central banking, legal, and policy perspectives

Considerations ⚠️

  • Author attribution — 99.16% of articles are tagged “Unknown,” which is historically accurate for unsigned editorials but limits author-based retrieval
  • Text length distribution — 44.5% of records are 100–500 characters (likely individual paragraphs or short notices); these are suitable for RAG but may need concatenation for fine-tuning contexts
  • Some OCR artifacts in titles — random sampling shows a minority of titles are sentence fragments rather than formal headings (a known OCR reconstruction challenge)
  • No explicit date column — temporal data is embedded within the issue identifier string (e.g., 1920-sim_economist_1920-03-06), requiring parsing for time-based queries

Topic Coverage Heatmap

Domain TopicArticles Mentioning% of Corpus
Stocks & Bonds597,96030.0%
Banking250,77112.6%
Credit95,2764.8%
Federal Reserve / Central Bank20,9501.1%
Interest & Discount Rates7,1440.4%
Inflation / Deflation4,8480.2%

RAG Readiness Assessment

CriterionStatus
Text chunking appropriate for embedding✅ Median 571 chars — ideal for vector search
Unique identifiers for source tracking✅ Issue field provides full provenance
Content-metadata separation✅ Article vs. metadata type flag
Semantic searchability✅ Natural language paragraphs, not tabular fragments
Citation traceability✅ Issue IDs encode publication + date

Recommendations for Optimization

  1. Parse issue identifiers into separate year, publication, and date columns to enable temporal filtering in RAG pipelines
  2. Flag and optionally merge very short records (<300 chars) with adjacent paragraphs from the same issue for richer context windows
  3. Add topic tags via embedding-based classification to enhance filtered retrieval across the 2M records

3 Sample Cortex Agent Prompts:

Here are three ready-to-use prompts demonstrating the dataset’s capabilities:

Prompt 1: Historical Monetary Policy Research

“What were the key discussions about Federal Reserve discount rate policy in the Commercial and Financial Chronicle during 1919–1921?”

This leverages the dataset’s strength in central banking coverage during the critical post-WWI tightening cycle that preceded the 1920–21 recession.

Prompt 2: Comparative International Finance

“How did The Economist report on British banking conditions compared to American credit markets during World War I (1914–1918)?”

This taps into the cross-publication diversity, enabling comparative analysis between the UK and US financial perspectives during a pivotal era.

Prompt 3: Trust & Fiduciary Law Evolution

“What were the emerging trends in trust administration and estate management practices discussed in Trusts & Estates magazine during the 1920s?”

This demonstrates the dataset’s niche coverage of fiduciary topics—an area rarely found in modern training corpora—making it valuable for specialized legal-financial AI applications.

Questions?

If you have questions about this dataset, or would like to speak about licensing, please refer to the Snowflake marketplace listing or contact us directly.

Citation

©Devin Media Corp. (2026). Finance AI: Historical Banking & Credit Archive (pre-1930). Available via Snowflake Marketplace.

Contact

hello(at)devinmediacorp.com

Scroll to Top