r/datasets 8h ago

dataset An Archive of 244 New York City Rent-Stabilization Histories (42 years) [JSON]

Thumbnail johnl.ink
5 Upvotes

42 years. 244 apartments. Four decades of rent, vacancies, improvements, preferential leases, and deregulations, as filed with the New York State Division of Housing and Community Renewal.


r/datasets 9h ago

request Open dataset for firm-level AI workforce / AI skills data?

3 Upvotes

Looking for an open-access dataset with firm-level AI workforce or AI skills data — global coverage, up to the present.

I've looked at Revelio Labs and Cognism, but all of them are paid licenses.

Is there anything open or free for academic use?

Thanks.


r/datasets 14h ago

dataset I trained a 67M-param LaTeX OCR model that runs on a laptop CPU — and built a new style-aware dataset to train it. Weights, data, and training code all open (MIT).

2 Upvotes

Hey everyone! I've been working on a little side project I wanted to share: latex-ocr, a standalone formula OCR model — you feed it an image of a math formula, it spits out the LaTeX source.

The main hook: it's only 67M parameters, so it runs comfortably on a laptop CPU. No GPU, no 300M-parameter monster to load. It's a CoCa-style model (contrastive captioner adapted for OCR), and despite the small size it beats the 107M UniMER-tiny baseline and gets pretty close to the 325M one on plain formulas.

The part I'm actually most proud of is the dataset. Real papers don't just use plain symbols — you see \mathbb{R}, \mathcal{F}, \mathfrak{g} everywhere, and existing OCR datasets basically ignore font styles, so models trained on them can't read (or hallucinate) those macros. So I rebuilt ~1.3M formulas with a MathJax → SVG → PDF → PNG pipeline and injected font-style macros with semantic heuristics (number sets → \mathbb, vectors → \mathbf, differentials → \mathrm). On that styled test set it clearly outperforms all the baselines — fair warning though, those baselines are zero-shot on styled data, so take that comparison with a grain of salt. The plain-split numbers are the like-for-like ones.

Everything is open: model weights and dataset on Hugging Face, training recipes included if you want to reproduce or fine-tune it yourself, MIT license. There's also a FastAPI server and a Gradio web UI, so you can drag-and-drop an image and see the LaTeX with a rendered preview.

Repo: https://github.com/PadishahIII/latex-ocr Model: https://huggingface.co/PadishahIIIXXX/latex-ocr Dataset: https://huggingface.co/datasets/PadishahIIIXXX/latex-ocr-dataset

Happy to answer questions about the training setup, the data pipeline, or anything else. Would love feedback — especially if you try it on your own gnarly formulas and it breaks, that's genuinely useful.


r/datasets 3h ago

API Title: Looking for one pilot user for a normalized SEC Forms 3/4/5 API

1 Upvotes

I’ve been building a normalized dataset and API from SEC ownership filings (Forms 3, 4 and 5, including amendments). It currently covers roughly 869,000 filings and more than 2.4 million transaction line items, including issuer and reporting-owner data, roles, transaction codes, timestamps, and price-quality information.

I’m not launching it as a public API yet. I’m looking for one person who already works with SEC ownership data and has a concrete research, backtesting, or data-engineering task they would be willing to run against it.

The pilot currently provides one filtered transaction endpoint with cursor pagination through an individual, revocable API key. There is no bulk export or separate filing/footnote endpoint yet.

The documentation is available here:

https://insiderfilings.info/api/v1/docs/

If this fits something you are already working on, please comment or DM me with a sentence or two about your use case. In return for access, I’m mainly looking for candid feedback on whether the API actually saves work, where the data model is unclear, and what blocks the workflow.


r/datasets 6h ago

request Egocentric POV / OTS Data Needed, Large Number of Hours | Worldwide

1 Upvotes

We're an AI training data marketplace working with companies building robotics and physical AI systems. We're currently looking for a data supplier, collector, or lab that has a large amount of egocentric video data and is willing to sell or license it.

Currently looking for 10,000 to 50,000 hours, with priority for OTS (off-the-shelf), already recorded material. This could be previously sold or unsold. We're interested in various tasks and regions, in both residential and commercial settings.

We have the budget to purchase, and the timeline is 3-6 months.

If you're able to produce those hours, let us know what your collection volumes per month look like. Happy to consider if your data and price are the right fit.

We're also interested in various setups (IMU, stereo, hand tracking, annotation, narration, etc.).


r/datasets 8h ago

request Anyone from institutions which provide access of dataful.in, we need data of MPLADS for our SIH project, so if there's anyone who can help...

Thumbnail
1 Upvotes

Doo heelppp guyssss plzzzz 😭😭


r/datasets 10h ago

resource Built a qualitative analysis tool — offering free sentiment/thematic analysis to test it on real data

1 Upvotes

I built a tool for qualitative research and I need people to break it.

I’m a student with a background in ML/data science, and I’ve been building a tool that can analyze things like interview transcripts, focus groups, open-ended survey responses, reviews, etc. — basically qualitative text.

It does sentiment analysis, coding, and thematic analysis.

But here’s the problem: I don’t want to test it on fake ChatGPT-generated data.

I’m looking for researchers/students who have real qualitative data they’ve already collected and wouldn’t mind letting me run through the tool.

In return, I’ll send you the analysis/results completely free.

No subscription. No sales pitch. No “book a demo.” 😅

I’m mainly trying to figure out:

\- Are the codes actually useful?

\- Do the themes make sense?

\- Does it save you time?

\- Where does it completely screw up?

\- Would you actually use something like this in your research?

You can anonymize/redact anything sensitive before sending it.

If you have a thesis, dissertation, research project, interview transcripts, focus group data, survey responses, or anything else qualitative sitting on your laptop, I’d genuinely love to test it.

Comment “interested” or DM me.

And yes, I’m specifically looking for people who are willing to tell me “this is terrible” if it is. 😂


r/datasets 10h ago

API [Self-promo: I built it] Real SEC ownership-filing metadata (Form 4 / 13F) as a free JSON API — sample dataset of 23 tickers

1 Upvotes

Disclaimer per sub rules: I built and operate this — it is self-promotion, disclosed up front.

The dataset: real SEC EDGAR ownership-filing metadata — Form 4 (insider trades), 13F-HR (institutional holdings), Form 3 — for a sample of 23 major US tickers, last ~120 days (675 records currently). Original source is the SEC's public EDGAR system (public domain): https://www.sec.gov

Fields per record: ticker, CIK, form type, filing date, accession number, primary document, and the SEC source URL.

Access: free demo API key at https://edgarfeed.onrender.com (no card).

Honest limits: metadata-level records only (not parsed line items yet); sample coverage, not the full market; this is a beta demand test, not a finished product. If you need full history, webhooks, or normalized transaction line items, that's the planned paid roadmap ($0/$29/$79 — not yet verified).

Questions about the data or the approach are welcome.


r/datasets 14h ago

dataset [self-promotion] Daily dataset + free API for Pakistani mutual fund NAVs (MUFAP-sourced)

1 Upvotes

Disclaimer: this is my own open-source project.

I maintain a daily-refreshed dataset of Pakistani mutual fund NAVs, built from the public MUFAP data (the industry association at mufap.com.pk is the original source). It is served as a free no-key API.

Per fund you get the current NAV, full NAV history and computed returns, plus AMC and category metadata. JSON, MIT licensed, no PII.

https://github.com/saadsalmankhan/pakistan-mutual-funds-api

Happy to expose extra fields if they are useful for research.


r/datasets 22h ago

mock dataset Free High-Fidelity Dataset - Retail POS Transaction

Thumbnail github.com
1 Upvotes

A high-fidelity synthetic retail POS (Point of Sale) transaction dataset containing over 1,190,000+ synchronized master invoices and 4,190,000+ item line details across multiple relational schemas. Generated using a custom, highly-optimized Prolog simulation engine, this dataset is enterprise production-grade and perfect for database stress-testing, query optimization benchmarks, and complex retail machine learning models.


r/datasets 9h ago

request Hi , I’m Pia : urgently looking for a data based role

Thumbnail
0 Upvotes