Finance AI: Historical Banking & Credit Archive (pre‑1930)
Dataset Documentation
Overview:
This dataset is a substantial, well-structured historical financial corpus comprising over 2 million records drawn from 10 authoritative financial publications spanning from 1843 to the early 1930s. It has undergone rigorous OCR processing and paragraph reconstruction, resulting in excellent core data completeness.
AI Training Readiness Report:
*Overall AI Training Readiness Score: 84 / 100
Dataset readiness report prepared by Dataset IQ.
Dataset Profile
| Metric | Value |
|---|---|
| Total Records | 2,016,529 |
| Article Records | 1,992,593 |
| Metadata Records | 23,936 |
| Distinct Publications | 10 |
| Distinct Issues | 7,779 |
| Temporal Span | ~1843 – 1930s |
| Avg. Text Length | 846 characters |
| Median Text Length | 571 characters |
Source Publications
The dataset aggregates content from 10 high-quality historical financial periodicals:
- The Economist (UK — global finance & trade)
- Commercial & Financial Chronicle (US — capital markets)
- Federal Reserve Bank of New York Monthly Review (US monetary policy)
- Citibank Monthly Economic Letter (commercial banking commentary)
- Lloyd’s Bank Review (UK banking perspective)
- Trusts & Estates (fiduciary & trust administration)
- Business Credit (commercial lending)
- Business Month (general business economics)
- Taxes (fiscal policy & taxation)
- The Annalist: A Magazine of Finance & Commerce (financial analysis)
Scoring Breakdown
| Category | Score | Notes |
|---|---|---|
| Volume & Scale | 19/20 | 2M+ records is exceptional for domain-specific AI training |
| Data Completeness | 16/20 | 0% null text, 0% null issues; author field is 99% “Unknown” (expected for historical periodicals) |
| Text Quality & RAG Readiness | 16/20 | Pre-chunked paragraphs at ideal retrieval lengths; no empty texts; min length 201 chars ensures meaningful content |
| Structural Consistency | 17/20 | Clean schema with type separation, consistent issue identifiers encoding year + publication + date |
| Domain Richness | 16/20 | 10 sources across banking, securities, monetary policy, taxation, trust law; 86+ year temporal breadth |
Strengths ✅
- Zero null/empty text fields — every record contains substantive content
- Pre-chunked paragraph structure — the median length of 571 characters is ideal for RAG retrieval without further splitting
- Rich domain vocabulary — 30% of articles reference stocks/bonds, 12.6% discuss banking, 4.8% cover credit markets
- Temporal depth — spans the formation of the Federal Reserve (1913), WWI-era finance, and the lead-up to the Great Depression
- Publication diversity — covers US, UK, commercial, central banking, legal, and policy perspectives
Considerations ⚠️
- Author attribution — 99.16% of articles are tagged “Unknown,” which is historically accurate for unsigned editorials but limits author-based retrieval
- Text length distribution — 44.5% of records are 100–500 characters (likely individual paragraphs or short notices); these are suitable for RAG but may need concatenation for fine-tuning contexts
- Some OCR artifacts in titles — random sampling shows a minority of titles are sentence fragments rather than formal headings (a known OCR reconstruction challenge)
- No explicit date column — temporal data is embedded within the issue identifier string (e.g.,
1920-sim_economist_1920-03-06), requiring parsing for time-based queries
Topic Coverage Heatmap
| Domain Topic | Articles Mentioning | % of Corpus |
|---|---|---|
| Stocks & Bonds | 597,960 | 30.0% |
| Banking | 250,771 | 12.6% |
| Credit | 95,276 | 4.8% |
| Federal Reserve / Central Bank | 20,950 | 1.1% |
| Interest & Discount Rates | 7,144 | 0.4% |
| Inflation / Deflation | 4,848 | 0.2% |
RAG Readiness Assessment
| Criterion | Status |
|---|---|
| Text chunking appropriate for embedding | ✅ Median 571 chars — ideal for vector search |
| Unique identifiers for source tracking | ✅ Issue field provides full provenance |
| Content-metadata separation | ✅ Article vs. metadata type flag |
| Semantic searchability | ✅ Natural language paragraphs, not tabular fragments |
| Citation traceability | ✅ Issue IDs encode publication + date |
Recommendations for Optimization
- Parse issue identifiers into separate year, publication, and date columns to enable temporal filtering in RAG pipelines
- Flag and optionally merge very short records (<300 chars) with adjacent paragraphs from the same issue for richer context windows
- Add topic tags via embedding-based classification to enhance filtered retrieval across the 2M records
3 Sample Cortex Agent Prompts:
Here are three ready-to-use prompts demonstrating the dataset’s capabilities:
Prompt 1: Historical Monetary Policy Research
“What were the key discussions about Federal Reserve discount rate policy in the Commercial and Financial Chronicle during 1919–1921?”
This leverages the dataset’s strength in central banking coverage during the critical post-WWI tightening cycle that preceded the 1920–21 recession.
Prompt 2: Comparative International Finance
“How did The Economist report on British banking conditions compared to American credit markets during World War I (1914–1918)?”
This taps into the cross-publication diversity, enabling comparative analysis between the UK and US financial perspectives during a pivotal era.
Prompt 3: Trust & Fiduciary Law Evolution
“What were the emerging trends in trust administration and estate management practices discussed in Trusts & Estates magazine during the 1920s?”
This demonstrates the dataset’s niche coverage of fiduciary topics—an area rarely found in modern training corpora—making it valuable for specialized legal-financial AI applications.
Questions?
If you have questions about this dataset, or would like to speak about licensing, please refer to the Snowflake marketplace listing or contact us directly.
Citation
©Devin Media Corp. (2026). Finance AI: Historical Banking & Credit Archive (pre-1930). Available via Snowflake Marketplace.
Contact
hello(at)devinmediacorp.com