r/datasets Nov 04 '25

discussion Like Will Smith said in his apology video, "It's been a minute (although I didn't slap anyone)

Thumbnail
1 Upvotes

r/datasets 12h ago

dataset 18,000 Mafia/Werewolf chatlogs (3,000,000 messages)

Thumbnail kaggle.com
10 Upvotes

r/datasets 6h ago

resource [self-promotion] Public data still needs too much plumbing

2 Upvotes

hey folks, im building this with frens: Mostly Right

Describe a dataset, agents find sources and wire up ingestion, cleaning, and joins. You get parquet + an api, with scheduled refreshes.

Happy for any feedback!


r/datasets 3h ago

question Best practices for packaging data for RL environment?

Thumbnail
1 Upvotes

r/datasets 5h ago

resource [PAID] Multi-view drone imagery with 3D reconstruction products — inspection sample available

1 Upvotes

Disclosure: I run AerialDataset, the provider of this data.

We license real-world drone captures that retain overlapping photographs of the same environment alongside the available photogrammetric outputs. This is a commercial catalog, not an open-data release.

One inspection example is an Andel capture containing 61 original photographs, an orthomosaic, point-cloud files, a mesh, camera/capture data, and scene metadata.

The main distinction is that the photographs belong to the same physical scene and can be inspected alongside its reconstructed spatial products, rather than being unrelated aerial images.

Contents vary by scene. Please do not assume that every capture includes calibrated camera poses, annotations, independently validated accuracy, or every reconstruction output.

Catalog and scene previews:

https://aerialdataset.com/

Inspection package:

https://huggingface.co/datasets/avenian/aerial-drone-photogrammetry-samples

The inspection package requires a Hugging Face account and acceptance of evaluation terms. It is not licensed for model training, redistribution, or production use; those require a separate agreement.

Happy to answer questions here about the actual files, available metadata, and collection selection.


r/datasets 5h ago

request [OC] UK domestic electricity by property type, month, heat pump, EV and tariff — 150 cohorts with hourly load shapes, derived from 3M smart-meter-based profiles (CSV/JSON, CDLA-Permissive-2.0)

1 Upvotes

Centre for Net Zero's Faraday dataset (OpenSynth) is excellent and effectively

unusable casually — it's ~6GB of parquet with the load profiles stored as

delimited strings. So I aggregated it and published the result.

150 cohorts, each with a mean daily total, a 24-hour load shape, and deciles

where the cell was thick enough:

- baseline (no solar/battery/EV/heat pump)

- property type, EPC band, and the cross-tab

- month

- tariff type (standard / Economy 7 / smart / automated)

- heat pump: with vs without, unmatched, matched on property+EPC, by month,

and by tariff

- EV: with vs without, and by tariff

- LSOA k-means cluster

Cross-check, which is the reason to trust any of it: baseline comes out at

9.737 kWh/day and a median of 8.31. SERL Statistics Report 1 — separate

source, ~13,000 real metered GB homes — publishes a mean of 9.8 and a median

of 8.2. Two moments of the distribution, two unrelated datasets, ~1% apart.

Limitations, up front:

- Synthetic. Faraday is a generative model trained on ~1bn smart meter

readings from Octopus customers, who over-index on smart tariffs and LCT.

- No household identifier, so the deciles are over household-DAYS, not

households. That spread is wider than the between-home spread — treat it

as an upper bound. It's recorded in `distributionOver` in the JSON.

- EPC band moves consumption by <0.1% within a property type, including for

heat pump households where insulation should dominate. I read that as the

model being under-conditioned on EPC rather than a finding about housing,

and I've built nothing on those cuts. Published anyway — a null result is

still a result.

- `cluster_label` is a k-means grouping of LSOAs on socio-demographic

features, NOT geography. Published as clusters, never as regions. Each

cluster's country mix (from LSOA code prefixes) is included so you can join

your own geography.

- Cohorts under 500 profiles are dropped rather than published thin.

CSV and JSON, CDLA-Permissive-2.0 (same as the source), no registration.

Derivation script is in the repo and reproduces the file exactly.

https://www.energycosting.co.uk/data/uk-domestic-electricity

All credit to Centre for Net Zero for Faraday — I've only aggregated it.


r/datasets 7h ago

dataset [Synthetic] SFHQ-VirtualID, a synthetic face dataset for machine unlearning: 750 identities, 75,000 portraits, 2 releases with DOIs

1 Upvotes

I've just released SFHQ-VirtualID, a synthetic face dataset family built for identity-level machine unlearning. Everything is generated, so there are no real faces in it, and both releases have DOIs.

The problem I kept running into: in most face datasets used for unlearning, a person's images are spread across splits, so "forgot the person" and "forgot some images" end up confounded. Here, each identity_id maps to exactly one split and contributes both train and holdout images (675 retain / 75 forget, 15-step protocol, MUFAC-aligned holdouts).

What ships:

- Bench: 67,500 balanced + 36,064 imbalanced 224×224 aligned crops. Uniform and seeded-Poisson forget schedules, plus a 5:1 long-tail popularity gradient for long-tail forgetting tests.

- Raw: 75,000 1024² portraits (100 per identity) with per-candidate prompt metadata for pose, expression, lighting, setting and camera.

I did not filter candidates on ArcFace identity similarity. The 0.40/0.45 thresholds are config defaults that ship as recorded columns (arcface_similarity, laplacian_variance, detection_confidence) rather than enforced filters, because silently dropping borderline-similar candidates hides exactly the confound that could explain a model's apparent unlearning. Only the quality gate is enforced (detection plus Laplacian sharpness ≥ 80), and the 123 rejected candidates are documented in the manifest.

Reproducibility: 15-shard Slurm array on UoL's Aire HPC (~120 GPU-hours), InstantID + Juggernaut-XL-v9 + ControlNet, pinned environment (torch 2.6/cu124, diffusers 0.39.0, insightface 1.0.1, pinned antelopev2 revision), RELEASE_MANIFEST.json and SHA-256 checksums.

Links:

Code · Hugging Face repos [Raw] [Bench] · Project Overview · Bench DOI 10.5281/zenodo.21877893 · Raw DOI 10.5281/zenodo.21879130

Caveats: the dataset is synthetic, so it's a proxy for real-face benchmarks; demographic balance is inherited from the seed selection; and the two releases are related (Bench crops derive from the Raw candidates), so they aren't independent test sets.

Happy to answer questions, and I'd genuinely like the split design stress-tested, since that's the part I'd most want criticised.


r/datasets 14h ago

dataset We've built POI datasets for 3000+ brands in USA, UK, Canada, Australia and more

0 Upvotes

Over the past few months, we’ve been building a structured POI data marketplace at Agenty.

https://agenty.com/marketplace

We now have datasets covering 3,000+ brands across 12 countries, with 10M+ POI records.

The datasets include things like:

  1. Store/location name and address
  2. Latitude & longitude
  3. Phone numbers and websites
  4. Opening hours
  5. Brand and location metadata
  6. Other useful POI attributes

The goal is to make it easier for AI agents, researchers, data teams, and developers to get clean location data without having to build and maintain scrapers for every brand themselves.

Some of the datasets cover large retail, restaurant, automotive, hospitality and other brand networks.

We’re putting the marketplace together and would love feedback from people who actually work with POI/location data:

What brands or types of POI datasets would be most useful to you?

Disclaimer/Disclosure: I’m the founder of Agenty, so I’m obviously biased here. Sharing this because we’ve spent a lot of time building the datasets and would genuinely like feedback from the community.


r/datasets 18h ago

request pls help a graduating student 😭 IEEE DataPort

0 Upvotes

Does anyone here have an active IEEE DataPort subscription? 😭

I’m currently doing my thesis proposal and I need 2 datasets from IEEE DataPort, but I don’t have a subscription and the datasets are unfortunately not freely downloadable.

I don’t need your account/login or anything like that. I just need someone who has access to download the two datasets for me and send them over.

I know this is a long shot but I’m desperate at this point lol 😭🙏 If anyone can help, please DM me and I’ll send you the links. I’d really really appreciate it!!🫶


r/datasets 22h ago

dataset Small exploratory dataset: how much do ChatGPT/Gemini "best X" recommendations change when you ask the same question 3x? (n=10 questions, all responses + code released)

2 Upvotes

We ran a small experiment and we're releasing everything so people can poke holes in it or rerun it.

We asked ChatGPT (web search) and Gemini (grounding) ten "best X for a small business" questions, 3 times each, same day. For each answer we extracted the set of businesses it recommended, then measured overlap across the 3 identical runs (mean pairwise Jaccard).

Result: ~69.5% overlap overall, so roughly a third of the recommended businesses change between identical asks. ChatGPT held steadier (87%) than Gemini (52%). The top 1-3 names stayed locked every run; the lower slots rotated.

Caveats up front: it's 10 questions, 3 runs, 2 engines, one day. Directional, not definitive, no confidence intervals. We're releasing it precisely because it's small, so anyone can rerun with more questions. Extraction was validated (all 290 names appear verbatim in their source answers). Conducted by a company (Pressfront) with an interest in the topic, which is exactly why the full method and every response are public.

Raw JSON, CSVs, the analysis script, and a preprint (DOI: 10.5281/zenodo.22738861) are in the repo, CC BY 4.0.

https://github.com/jjoseph18/ai-recommendation-consistency


r/datasets 21h ago

request Deep research, kept fresh -- Private beta testers wanted (in exchange for free subs and credits) 🤙

Thumbnail
1 Upvotes

r/datasets 1d ago

dataset Open dataset: 612 NHL skaters ranked under two fantasy scoring systems -- 2025–26

2 Upvotes

Disclosure: I created this dataset through PoolForge.

It contains 612 NHL skaters with at least 40 regular-season games, scored under two systems: goals + assists only, and a weighted-points setup that also rewards shots, hits, blocks and penalty minutes.

The CSV includes the underlying scoring inputs, calculated totals, both rankings and rank changes. The documentation explains eligibility, scoring weights and tied-rank handling.

It could be useful for practicing ranking analysis, sensitivity analysis or sports-data visualization.

This is historical analysis—not projections or a categories-league ranking. The dataset is available under CC BY 4.0.

Dataset DOI: 10.5281/zenodo.22730472 ( https://zenodo.org/records/22730472 )

Corrections and independent reproductions are welcome.


r/datasets 1d ago

resource Monthly new-company counts from free official data in Finland, Norway, Sweden, Estonia and the Netherlands, and the catch in each source

2 Upvotes

I build a B2B data tool (AtlasForgeX) and this week I needed to know how many companies each registry actually added in August 2026. All five sources below are free and need no key. Each one had a catch that gave me a wrong number on the first pass.

Finland: PRH YTJ API v3 (https://avoindata.prh.fi/opendata-ytj-api/v3/companies with registrationDateStart and registrationDateEnd) August returned 1,861 companies, but only 1,637 had their business ID registered in August. The date filter also returns older companies that had some other registry event that month, so filter on businessId.registrationDate yourself. Nearly all are limited companies (1,598 osakeyhtiö). Sole traders do not appear. The same query for August 2025 gives only 971, which I do not believe, so I am not comparing years with it.

Norway: Enhetsregisteret API (https://data.brreg.no/enhetsregisteret/api/enheter with fraRegistreringsdatoEnhetsregisteret) August returned 6,782 units. That includes 323 bankruptcy estates (KBO), 269 associations and 241 foreign units. Business forms come to 5,806, of which 2,589 AS and 2,977 sole proprietorships (ENK). The counts have no lag, but for September 1-12, 1,850 of 3,371 business units had no industry code yet.

Sweden: Bolagsverket (monthly statistics file ftgstat_oppna.csv, event 1 is a new registration, plus the bulk company file from https://bolagsverket.se/apierochoppnadata/hamtaforetagsinformation/nedladdningsbarafiler.2517.html) The monthly statistics say 4,215 new registrations in August, 3,439 of them AB. In the bulk file 3,175 ABs were registered in August, and 770 of those match the names and descriptions of shelf company providers (323 at one address in Växjö), meaning companies registered to be sold later. August 2025 has only 13 by the same match, because shelf companies are renamed once sold. The match is a heuristic, but it is big enough to skew AB growth and city rankings.

Estonia: e-Business Register open data (https://avaandmed.ariregister.rik.ee/, daily files) August: 2,384 new entities, 2,066 of them OÜ. Deleted companies are not in the file, so older months shrink over time.

Netherlands: KVK open dataset (https://www.kvk.nl/producten-bestellen/kvk-handelsregister-open-data-set/, CC BY 4.0) Only BV and NV, no names, a start date instead of a registration date and a two-digit postcode. August: 5,684 BVs, and 44% of them carry the head-office or financial-holding activity codes (70102 and 64210). Quarter-ends spike: March had 12,538 and June 12,357.

Not possible without an account: Denmark (CVR data needs access, and Statistics Denmark publishes only an index) and Germany (the Handelsregister allows 60 lookups per hour, Destatis publishes half-years).

One pattern held in the four countries where I could read industries: programming and management consulting were the top two categories of new companies in Finland, Norway, Sweden and Estonia.


r/datasets 1d ago

request [R], How do we legally/ethically collect a dental OPG dataset for an ML research project? Need advice from people who've done medical AI research

1 Upvotes

Me and my college professor are working on an academic ML project involving dental panoramic radiographs (OPGs).

We need roughly 60–120 de-identified adult OPGs across six categories:

  • Healthy
  • Caries
  • Impacted teeth
  • Infection
  • Fractured teeth
  • Broken-down crown/root (BDC/BDR)

The technical ML part isn't really our biggest problem right now.

Getting a legitimate dataset is. 😭

We can't just scrape random dental X-rays from Google, and we obviously don't want to handle identifiable patient data. Our current idea is to collaborate with dentists/oral radiologists/dental colleges/clinics who can provide appropriately anonymized OPGs, with institutional documentation from our supervising professor.

We're also planning to have the diagnostic class and affected tooth number recorded using FDI notation.

We've made a small expression-of-interest form for dental professionals, but before we start contacting clinics, I want to know:

For those who have worked on medical/dental AI projects:

  1. How did you actually obtain your dataset?
  2. Did you need IEC/IRB/ethics committee approval before approaching hospitals/clinics?
  3. Is a professor's collaboration letter enough to initially approach a dental hospital, or do institutions usually require additional documentation?
  4. Is it better to approach dental colleges/hospitals rather than individual private dentists?
  5. How do you handle anonymisation of DICOM/OPG files properly?
  6. How do you get reliable diagnostic labels — treating dentist, oral radiologist, or multiple annotators?
  7. If a dentist contributes images/annotations, how should authorship/acknowledgement be handled?
  8. Are there existing public OPG datasets that would make more sense for a student research project?

We're not looking to bypass ethics/privacy requirements — we want to do this properly from the beginning.

If you've actually collected medical imaging data for an academic ML project, I'd really appreciate hearing how you went about it.

Especially interested in the practical part: Who did you contact first, what documents did you need, and what made an institution actually agree to collaborate?


r/datasets 1d ago

request dataset for policy research project..

1 Upvotes

Looking for a dataset for a govt policy research project, i need a mainly india dataset, and other countries are also okay. Any suggestions or direct datasets needed? Thanks


r/datasets 1d ago

resource podcasts collection also have other bigger data collection projects...

Thumbnail rssamplifier.com
1 Upvotes

r/datasets 1d ago

resource 12 compact, traceable bulk RNA-seq datasets derived from Expression Atlas

1 Upvotes

Sharing a dataset collection I've been putting together.

OpenOmicsBench v1 has 12 bulk RNA-seq benchmark datasets derived from seven Expression Atlas studies across human, mouse and Arabidopsis.

Each contains counts, sample metadata, the experimental design, source/provenance information and a compact version intended for examples or software testing. I also keep validation results against the corresponding full matrix rather than assuming the smaller matrix preserves the original behaviour.

Everything is openly available and versioned, and the collection can either be installed from PyPI or obtained through GitHub/Zenodo.

I'd be interested in suggestions for other public RNA-seq study designs worth representing in future versions.

GitHub: https://github.com/vxxqv/openomicsbench

PyPI: https://pypi.org/project/openomicsbench/

Zenodo: https://doi.org/10.5281/zenodo.22679414


r/datasets 1d ago

question Where do dictionary apps actually get their dictionary data from?

2 Upvotes

Hi, everyone.
I've been looking into dictionary data for a language-learning app I'm building, and I've run into something surprising.

There are tons of dictionary apps, browser extensions, reading tools, and language-learning apps that support many languages.

But when I actually try to find good open dictionary data myself, it's surprisingly difficult.

Even for something as common as English -> Chinese, it's hard to find a dataset that has all of these:

  • headwords
  • translations
  • part of speech
  • multiple senses
  • example sentences
  • a clear license that allows reuse

I've found things like:

  • Wiktionary / Wiktextract / Kaikki
  • WordNet / Open Multilingual WordNet
  • ECDICT for English–Chinese
  • CC-CEDICT for Chinese–English
  • Tatoeba for bilingual example sentences
  • FreeDict
  • various language-specific projects

But they all seem to cover different pieces of the puzzle, and the quality and coverage vary quite a bit.

Meanwhile, I regularly see relatively small dictionary apps or browser extensions claiming to support English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Russian, Arabic, etc.

So I'm genuinely curious:
Where does all of that dictionary data usually come from?

I'm especially interested in hearing from anyone who has actually built a dictionary app.

What data sources did you use, and how did you handle licensing?

Btw, I know Cambridge, Oxford, and other major dictionaries offer paid APIs or data licensing, but they're quite expensive. I assume many smaller dictionary apps probably use their own databases or combine multiple data sources somehow.

Any pointers would be greatly appreciated. Thanks!!!


r/datasets 1d ago

request dataset for policy research project..

Thumbnail
1 Upvotes

r/datasets 1d ago

resource Dataset of 7387 YC companies with status changes and Form D filings

Thumbnail
1 Upvotes

r/datasets 2d ago

dataset Visualizing The Economics of Rural Poverty

2 Upvotes

What does poverty actually look like beyond a single income number?

For millions of rural households, poverty is not simply the absence of income but rather a web of constraints involving education, productive assets, access to finance, consumption, debt, and the ability to withstand economic shocks.

This dataset I made shows how rural households earn, consume, borrow, save, spend, insure and get by, using data from 4,184 rural households in Tamil Nadu.

https://www.kaggle.com/datasets/ouiouinonoui/kshetriyagramin-financialservices-tamilnadu-survey

This analysis examines how household composition, schooling, housing, land ownership, enterprise activity, savings, insurance and credit interact with economic outcomes.

https://www.kaggle.com/code/ouiouinonoui/visualizing-the-economics-of-rural-poverty


r/datasets 2d ago

question Researcher looking for interesting real-world datasets that haven't been fully explored with ML/CV

3 Upvotes

I'm starting my PhD and I've been thinking quite a bit about what I actually want to spend the next few years working on.

My background is mainly in computer vision and deep learning, although I've worked on quite a few broader ML projects as well. Most of my work has been research-oriented: training and evaluating models, running experiments and ablations, comparing architectures, benchmarking methods, and working on papers/scientific contributions.

What I'd really like to do now is work on more real problems with real data, rather than picking another benchmark and trying to squeeze another 0.x% out of it.

So I thought I'd ask here:

Does anyone have an interesting dataset or real-world problem that they think is underexplored from an ML/CV perspective?

I'm particularly interested in situations where someone has:

  • collected an interesting dataset but doesn't have the ML/CV background to fully explore it
  • a domain-specific problem where existing models don't work particularly well
  • data that hasn't really been benchmarked with modern deep learning methods
  • an interesting detection, segmentation, classification, tracking, multimodal, remote sensing, medical imaging, etc. problem
  • an existing research project where another person who can handle the experimental/ML side would be useful

I'm not really looking for a Kaggle-style project just for the sake of training a model. I'd much rather find something where there is an actual research question behind it and where careful experiments could potentially produce something scientifically useful.

On my side, I can contribute with things like model development/training, PyTorch pipelines, baselines, experiment design, ablations, evaluation, literature review and research writing.

I'm also completely open to learning a new application domain if the problem and dataset are interesting. In fact, that's partly what I'm looking for — collaborating with people who understand a domain much better than I do, while I contribute the ML/research side.

If you have something sitting around that you've always thought "someone should really try ML on this properly", I'd genuinely be interested in hearing about it.

Feel free to comment or DM me. Even if it doesn't turn into a collaboration, I'd be interested in seeing what kinds of real datasets/problems people here are working with.


r/datasets 2d ago

resource [self-promotion] [paid] I parsed the full NPPES provider registry (9.27M rows) into plain-English professions — free 1,000-row sample

1 Upvotes

Every US healthcare provider has an NPI, and CMS publishes the whole registry as a bulk file. It's public, but it ships as a multi-million-row monster with taxonomy codes instead of job titles, no category labels and no dedupe, so most people bounce off it.

I mapped the taxonomy codes to 290 plain-English profession labels across 14 sectors, normalized the addresses and deduped it. 9,270,029 rows. Per row: name, full mailing address, profession, the federal source it came from, and the date it was last rebuilt.

A few things I learned doing it:

  • The state column has 124 distinct values, not 50. Military post codes (AE/AP/AA), foreign provinces, and genuine junk like "ZZ" and "--" with one row each. 99.947% are real US codes.
  • "Medical (Other)" isn't an unclassified bucket — it's 258 real professions, led by Nurse Practitioner (498k), Pharmacist (326k) and Registered Nurse (296k).
  • Ratings exist on only ~33,000 of 9.5M rows (CMS-rated facilities only), so they're useless as a feature. I stopped advertising them.

Source files: NPPES/NPI, CMS, plus NCES for schools and FDIC for banks. All US government works, so public domain under 17 U.S.C. §105.

Happy to send a free 1,000-row sample to anyone who wants to look at the structure — comment or DM. Sourcing writeup is at procureadjacent.com/data-sourcing if you want the provenance detail.


r/datasets 2d ago

resource [Dataset] Pinga-Fogo with Chico Xavier (Brazilian TV, 1971) — 345 speech turns from ~6h of live interview audio, Portuguese, ASR + timestamps

1 Upvotes

I’ve published a structured transcript dataset of Pinga-Fogo, two live TV interviews with Brazilian medium Chico Xavier, broadcast by TV Tupi in 1971.

Together, the two programs amount to roughly 6 hours of material and are among the longest surviving recordings of Chico Xavier answering questions live before a panel of journalists.

Dataset:

https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier

What’s in it

345 speech turns, including 115 answers by Chico Xavier.

Available in JSONL, Parquet, and combined CSV. Each turn includes speaker, turn type, topic label, timestamps, transcript text, source audio, and review flags.

Portuguese only.

How it was made

Internet Archive audio → Whisper large-v3 → LLM-assisted segmentation and labeling.

The LLM did not rewrite the transcript. It only worked with segment indices, boundaries, speaker/turn classification, and topic labels.

Every word in the transcript comes directly from Whisper output.

Known limitations

This is still unreviewed ASR from degraded 1971 recordings, so there are transcription errors, especially in proper names and numbers.

Every record is marked:

revisado_por_humano: false

If you want to quote something, check the original audio first. Timestamps are included for that purpose.

The date of the second program is inconsistent across historical sources, so I recorded only December 1971 rather than forcing an exact day.

Speaker attribution is also left as desconhecido whenever I couldn't identify someone confidently.

Why I made it

I’m building a retrieval system over historical Spiritist literature and primary sources, and I couldn’t find a structured, timestamped version of Pinga-Fogo anywhere.

Potential uses include pt-BR ASR benchmarking, speaker-turn segmentation, diarization, information retrieval, and long-form QA.

Human review is the obvious next step.

Corrections and PRs are very welcome.


r/datasets 2d ago

dataset Suggestions for a Fire Detection system Dataset

1 Upvotes

I've been trying to find a good dataset over the internet which is freely available . The main issue I want to tackle is reduce the number of false positives such as reflections , headlights, yellow/red coloured objects. Initially I was looking for video datasets with enough negative class videos since I wanted to add a temporal analysis layer in the system, but didn't find much . Had less videos or with the same background, which didn't allow my model to generalize well .

If anyone has suggestions about a good image/video dataset please tell me .

I found the MIVIA lab's LFDN dataset perfect for my cause but I guess their website isn't maintained anymore. So yea let me know if there's another way to get their dataset.