All micro-internships
Micro-internship brief ≈7h in one sittinglow-code

Merge two messy real-world lists into one clean dataset

The hardest part of real data work isn’t analysis — it’s that "Central Technical Institute" and "Central Tech. Inst." may be the same place. Merge two real overlapping lists with AI matching the near-duplicates — and you catching where it’s confidently wrong.

Data AnalysisSpreadsheet AnalysisData CleaningAnalytical ReasoningBusiness Communication

Posted by The AI Internships

The work brief

  1. 01Find two real overlapping lists: two college/coaching lists, two price lists, exam centre lists from different sites, or two years of the same report.
  2. 02Standardise both: consistent columns, formats, spellings. Use AI to help match near-duplicate names — but verify its matches; it will be confidently wrong.
  3. 03Merge them, marking for each row: matched, only-in-A, only-in-B, or conflict (and how you resolved it).
  4. 04Report the mess honestly: match rate, worst conflicts, and 2 things the merged data reveals that neither list showed alone.

What you’ll produce

4 deliverables

Submission standard

Submit both sources, the merged CSV with match-status column, and your merge report. Judgement calls are the point — document them. Include screenshots of both original lists (with their origin and date visible) so the merge can be checked against the real sources.

  • Your two sources: what they are, where from, how they overlap

    Give a link or a dated screenshot for each of the two sources: what they are, where from, and how they overlap.

    Written responseRequired
  • The merged dataset as CSV, including a match-status column

    Paste comma-separated data with one header row and one record per line. Use consistent column names and remove private information.

    CSV dataRequired
  • Merge report: match rate, hardest conflicts and your calls, where AI matching failed

    Written responseRequired
  • 2 things the merged data shows that neither source showed alone

    Written responseRequired

You’ll complete these inside your private workspace.

What you must submit as proof

This brief requires evidence an AI can’t fabricate.

  • Screenshots of BOTH original source lists (showing where each came from — site/header/date)

    Show the two real raw lists before merging so the merge can be checked against them.

    Image uploadRequired

Submissions without this evidence cannot be submitted.

Protect other people in your proof. Blur faces, names, phone numbers and email addresses before you upload, and refer to anyone you worked with by role or number ("Listener 1", "the stall owner"). Your proof is only ever used to check your work — it is never published, never appears on your certificate, and is never shown in your public portfolio.

How your work is evaluated

The passing benchmark is 70/100.

Merge quality

25%

The merged CSV is consistent, with honest match-status marking.

Documented judgement

25%

Conflicts and AI-matching failures are honestly documented with resolutions.

Merge payoff

13%

The insights genuinely require the combined data.

Real sources ↔ merged CSV

38%

The two source screenshots must be real, distinct lists that plausibly produce the merged CSV and its match-status marks. Invented sources or a merge that cannot be traced to them are a fail.

How we grade your AI usage

30% of your score

Using AI is the point — it’s the skill this certificate proves. You’ll answer three short questions about how you used it: what you asked, what was wrong with its first answer, and what you changed. Specific, honest answers score high. “I pasted the brief and submitted the answer” scores near zero.