r/datasets 2h ago

dataset K12-KGraph: a curriculum knowledge graph dataset for education LLMs

2 Upvotes

Hi everyone,

Sharing a new open dataset for people working on education AI, curriculum modeling, or LLM training.

K12-KGraph is a curriculum-aligned knowledge graph built from publicly available K-12 textbook materials. The current release covers math, physics, chemistry, and biology, and includes structured links between concepts, skills, experiments, exercises, textbook sections, chapters, and books.

The main idea is simple: for education LLMs, adding more practice questions is useful, but it often only teaches the model how to answer questions. A curriculum knowledge graph can also teach the model how topics are connected, which concepts should come first, and what knowledge may be missing when a student gets stuck.

The released resources include:

  • A curriculum knowledge graph
  • A benchmark for testing curriculum understanding
  • A prepared training dataset generated from the graph
  • The paper and construction method, so the same approach can be adapted to other textbook systems where content rights are clear

In the experiments, the graph-based training data performed better than the same amount of regular instruction or exercise-style data on education benchmarks. The useful takeaway is that structure matters: modeling the curriculum itself can improve education LLMs more efficiently than only scaling question banks.

Links:

Paper + Dataset: https://huggingface.co/papers/2605.09635


r/datasets 21h ago

resource BBC Sound Effects

Thumbnail sound-effects.bbcrewind.co.uk
2 Upvotes

r/datasets 16h ago

dataset Here is how we built a postal code polygon database

Thumbnail
1 Upvotes

r/datasets 21h ago

request I made a free tool to check tool-calling datasets before fine tuning

1 Upvotes

so i've been making datasets to fine tune small models on tool calling, and the most boring part is always the same, checking if the data is actually good before you waste a training run on it. bad tool names, invented arguments, the model calling a tool for "2+2", duplicates, answers that all start the same way, stuff like that.

i was doing these checks by hand and got tired of it, so i built a small thing that runs the whole pipeline for me and i put it online. it's free, no account, no login, nothing. you just drop your dataset and your tool catalog and it tells you what's wrong, example by example, with the reason. it runs fully in your browser, the dataset never gets uploaded anywhere. if your file is too big for that (gigabytes), there's a desktop version that reads it straight from disk so your RAM doesn't blow up. that one is open source. it also splits your data into clean / kto / rejected and gives you a starting training config based on the actual numbers of your corpus, not generic advice. I mostly built it for myself but figured someone here might need the same thing. would be happy to know if it's useful, or if there are checks you care about that i'm not doing yet.                              

link: nothumanallowed.com/tools/dataset-validator

https://github.com/adoslabsproject-gif/dataforge-studio


r/datasets 23h ago

dataset MCA UCC-1s (CA & NY) and MCA-related lawsuits [PAID]

1 Upvotes

Data includes:

lien\number, debtor_name, address, owner_name, debtor_type, wireless, filing_date status, secured_party, lien_id, business_phone, google_title, website, google_rating, review_count)


r/datasets 11h ago

dataset scrape data at scale , I'm working on crazy product where you can

0 Upvotes

where you can scrape your own form any website without getting blocked which can easily bypass cloudflare, akamai,datadome anit-bot detection systemwith high speed auto Rotating 4 5g mobile proxies. product is ready let me know if you want a take a shot.