Clean a messy dataset & find 3 insights
Real data is messy. Use an AI tool to clean a dataset, then surface three insights a decision-maker would care about.
Posted by The AI Internships
The work brief
- 01Take any messy CSV (or make one) with duplicates / blanks.
- 02Use AI (ChatGPT data analysis, a no-code tool, or code) to clean it — then spot-check its fixes against the raw rows; AI loves to "fix" data that was never broken.
- 03Ask AI what patterns it sees, verify the numbers behind its claims yourself, and write 3 insights a decision-maker would care about.
- 04Paste the cleaned CSV and attach one chart.
What you’ll produce
Submission standard
Paste your cleaned CSV (header + rows), your 3 insights, and a chart image — plus your AI workflow: tools, best prompts, and what you changed. Also paste the original messy CSV and a public link to your AI cleaning chat so the before→after and your process are both visible.
Your cleaned CSV (paste header + rows)
Paste header + rows of your cleaned CSV; it must plausibly be the cleaned version of the messy 'before' you also submit.
CSV dataRequired3 insights from the data
Written responseRequiredA chart of one insight
Image uploadRequired
You’ll complete these inside your private workspace.
What you must submit as proof
This brief requires evidence an AI can’t fabricate.
The original MESSY CSV (before cleaning) — same rows, duplicates/blanks still in
Graders diff this against your cleaned version; a pristine 'before' with nothing actually fixed is a fail.
CSV dataRequiredPublic share link to your AI cleaning chat (ChatGPT/Gemini 'Share')
Must be YOUR live chat showing the real cleaning back-and-forth, not a single dump.
Public linkRequired
Submissions without this evidence cannot be submitted.
Protect other people in your proof. Blur faces, names, phone numbers and email addresses before you upload, and refer to anyone you worked with by role or number ("Listener 1", "the stall owner"). Your proof is only ever used to check your work — it is never published, never appears on your certificate, and is never shown in your public portfolio.
How your work is evaluated
The passing benchmark is 70/100.
Data cleaning
29%The CSV is genuinely clean (consistent columns, no obvious blanks/dupes).
Insight quality
29%Insights are specific and supported by the data.
Visualisation
14%The chart communicates one insight clearly.
Before/after + real process
29%The messy 'before' CSV must plausibly clean into the 'after', and the shared chat must show real iterative cleaning. A polished CSV with no messy source and no process trace is a fail.
How we grade your AI usage
Using AI is the point — it’s the skill this certificate proves. You’ll answer three short questions about how you used it: what you asked, what was wrong with its first answer, and what you changed. Specific, honest answers score high. “I pasted the brief and submitted the answer” scores near zero.