Stress-test AI in your language
AI works great in English — but how good is it in Hindi, Marathi, or your language? Run 5 identical tasks in both languages, document exactly where it breaks, and report like a QA tester.
Posted by The AI Internships
The work brief
- 01Pick a language you actually speak (for example Hindi, Marathi, Spanish, Arabic, or another language used in your community).
- 02Design 5 tasks: translate a proverb, write a festival invitation, explain a school topic, handle a local place name, do word-play or grammar.
- 03Run each task through ChatGPT or Gemini in English and in your language — same tool, same day.
- 04Log where the non-English version is worse: wrong script, awkward phrasing, cultural mistakes, refused tasks.
What you’ll produce
Submission standard
Submit the language, all 5 side-by-side comparisons, a screenshot of the worst failure, and your findings. Add a public share link to the actual chat that shows the English and your-language runs side by side.
The language you tested
Short answerRequired5 side-by-side comparisons (prompt, English output, your-language output, verdict)
Written responseRequiredScreenshot of the worst failure you found
Screenshot must show the tool interface, your account, the non-English script rendered, and a visible date.
Image uploadRequiredYour findings: where does AI break in your language, and who does that affect?
Written responseRequired
You’ll complete these inside your private workspace.
What you must submit as proof
This brief requires evidence an AI can’t fabricate.
Public share link to the parallel-language runs
Your live ChatGPT/Gemini share link showing the same tasks run in English and in your language.
Public linkRequired
Submissions without this evidence cannot be submitted.
Protect other people in your proof. Blur faces, names, phone numbers and email addresses before you upload, and refer to anyone you worked with by role or number ("Listener 1", "the stall owner"). Your proof is only ever used to check your work — it is never published, never appears on your certificate, and is never shown in your public portfolio.
How your work is evaluated
The passing benchmark is 70/100.
Parallel testing
29%Same tasks genuinely run in both languages, with outputs shown.
Evidence
14%The screenshot shows a real failure a fluent speaker would catch.
Findings quality
29%Findings are specific and only writable by someone fluent in the language.
Verifiable parallel runs
29%Auto-fail if the link is missing/dead or doesn't show the same tasks run in both languages (invented comparisons).
How we grade your AI usage
Using AI is the point — it’s the skill this certificate proves. You’ll answer three short questions about how you used it: what you asked, what was wrong with its first answer, and what you changed. Specific, honest answers score high. “I pasted the brief and submitted the answer” scores near zero.