About
What this is
The Berkeley Stumper Challenge is an ArtifactQA competition for Berkeley students in consulting and finance. Your job is to author a realistic, hard task that requires research, reasoning, and a polished deliverable — a prompt that defines the work, a ChatGPT sample that shows where it falls short, structured task components, and a precise rubric that grades any response point by point.
Strong tasks ask for real expert work: a document, spreadsheet, slide deck, code file, or other artifact. They fix the date, name the audience, define the decision, set the sources, and include a quality bar high enough that even frontier models miss it. You don't need to code; you write in plain English in the domain you know cold. The real test is whether the task is genuinely hard and cleanly gradable.
Task anatomy
Every accepted submission contains four connected parts:
- 01
Final task prompt
A realistic expert request with role, audience, fixed date, specific scope, sources, deliverable format, and quality bar.
- 02
Sampled ChatGPT response
You run the prompt through ChatGPT, review the actual response and any artifact, and document the concrete failures that make it insufficient.
- 03
Task components
The same sampled prompt broken into structured fields: base prompt, target audience, persona, management preferences, reference context, and tooling.
- 04
Rubric
Binary, observable, weighted criteria — each tied to a prompt component — with sources and justifications so another reviewer can score any response consistently.
Why prompt–rubric pairs
Reinforcement learning needs a grader. Without one, a model can only be told that an answer was good or bad in aggregate — which teaches it style, not reasoning. A rubric turns your judgment into a reward signal: it names the exact steps a correct answer must contain, and the exact hedges and shortcuts that should cost points.
Most public evaluations in this space test recall of textbook definitions. Frontier models saturated those years ago. What they still fail is the working judgment of someone who has actually built the model, run the diligence, or sat in the deal room — where the answer depends on chaining several non-obvious steps and committing to one of them.
You design
A prompt and a weighted rubric from your domain.
We test
Frontier models attempt your puzzle; graders check the rubric.
Labs learn
Accepted puzzles become training signal for reasoning.
Example task
Below is an accepted submission, shown exactly as a reviewer sees it. Note how specific the rubric is: every line is something a grader can check, including a deduction for hedging. That specificity is what separates a strong puzzle from a weak one.
Prompt
“An acquirer pays 6.0x EBITDA at close plus an earn-out worth up to 1.5x EBITDA if the target hits its year-two plan. Walk through exactly how the earn-out flows through the acquirer's reported EPS in year two under US GAAP, explain why strong target performance can lower reported EPS, and name the single deal structure where that effect disappears.”
Rubric · 8 points
- +3Identifies that the earn-out is contingent consideration remeasured through P&L under ASC 805, not booked to goodwill after close.missed
- +2Notes the remeasurement makes forward EPS anti-correlated with target performance, and quantifies the direction.missed
- +2Distinguishes the accounting treatment from the cash-flow treatment in the DCF bridge.met
- +1States the one structure (equity-settled, fixed-share earn-out) where the effect disappears.missed
- −2Deduct if the answer hedges with 'it depends on the deal terms' without committing to a treatment.met
Domain coverage
We're targeting 250+ tasks across the areas below. Coverage is uneven today — credit, restructuring, and commercial diligence are the thinnest, so puzzles there are the most likely to stand out.
| Domain | Areas |
|---|---|
| Valuation & M&A | DCF & LBO modeling, Accretion/dilution, Purchase accounting, Synergy cases |
| Credit & Fixed Income | Covenant analysis, Spread decomposition, Restructuring waterfalls, Ratings logic |
| Markets & Trading | Derivatives payoffs, Market microstructure, Portfolio construction, Risk limits |
| Corporate Finance & FP&A | Three-statement logic, Working capital, Capital allocation, Scenario planning |
| Strategy Consulting | Market entry, Pricing strategy, Cost transformation, Competitive response |
| Operations & Diligence | Commercial due diligence, Supply chain economics, Post-merger integration, Unit economics |
Contribute
Who can contribute
Current Berkeley undergraduate and graduate students in consulting and finance — Haas, Economics, the MFE, and the campus consulting and investment clubs. What we're screening for is depth in a domain, not seniority: a student who has spent a summer in restructuring or two semesters on a diligence team can write a better puzzle than someone with a broad but shallow view.
Alumni and students at other schools can submit too, but the challenge's review capacity is reserved for Berkeley first through Aug 31, 2026.
What makes a strong task
Realistic and artifact-driven. Ask for a concrete deliverable — document, spreadsheet, slide deck, code, or PDF — with a named audience, decision, fixed date, and quality bar.
Research plus synthesis. Not just a list or lookup. Force comparison, calculation, diagnosis, recommendation, or judgment using conflicting evidence.
Specific and self-contained. Name the companies, cities, regulations, tools, and source hierarchy. A qualified human should be able to complete it without follow-up questions.
Tested against ChatGPT. Sample the prompt, review the actual response and artifact, and revise until the model shows concrete, explainable shortcomings. Don't make it hard with busywork.
Rubric-first. Criteria are binary, atomic, observable, prompt-aligned, weighted, and source-backed where factual. Every major requirement is covered exactly once.
Original. Not lifted from a casebook, prep guide, or existing benchmark. Models have already memorized those.
How to contribute
- 01
Pick your domain
Choose an area from the coverage table where you have real working experience. The thin areas — credit, restructuring, commercial diligence — are the most likely to stand out.
- 02
Draft a hard prompt
Set a role, audience, fixed date, specific scope, source hierarchy, deliverable format, and quality bar. Depth beats length; long prompts usually hide ambiguity.
- 03
Sample ChatGPT
Run the prompt in the submission platform's ChatGPT widget, sync the response and any artifact, and document concrete failures. Revise and resample until the model misses meaningfully.
- 04
Record validation notes and components
Break the sampled prompt into base task, audience, persona, management preferences, reference context, and tooling. These fields must match the prompt exactly.
- 05
Build the rubric
Write binary, atomic, observable, weighted criteria with sources and justifications. Run the autograder, fix real issues, then do a final manual pass.
- 06
Submit on the platform
Enter the task at qa-rubrics.vercel.app. You'll get reviewer feedback in the same place, usually within a week.
Review & scoring
Every submission goes through two passes. First a domain reviewer checks that the prompt is realistic and self-contained, the components match the sampled prompt, and the rubric is gradable — most winning tasks go through at least one round of edits. Then the task is run against a panel of frontier models and scored with your rubric.
A task scores well if it survives review and the models fail it cleanly. Before submitting, you should also run the built-in autograder and do a manual rubric pass: check coverage, binary scoring, artifact locations, source-backed factual claims, and no double counting. Submissions sourced from published material are rejected outright.
Prizes & Terac Fellows
The top performer wins $5,000. The top 20 overall contributors will be admitted as Terac Fellows, doing 10–40 hours per week of paid reasoning and research work — well compensated — alongside Terac and its lab partners. There is no separate compensation per individual task; the prize and fellowships are the reward for this challenge.
Fellows get a direct line to frontier-model research, written references from the Terac team, and first access to future paid projects. Contributors are credited by name in the dataset documentation unless you ask us not to be.
Deadline
Submissions close August 31, 2026. Review and iteration continue after that date, but no new tasks are accepted past it. Start early — nearly every strong puzzle needed a round of reviewer feedback first.
Resources
Links
- Submission platform — submit and track your tasks
- Rubric writing guide — worked examples and scoring conventions
- Slack community — ask questions and share drafts with other contributors
- Terac expert profile — how payment is handled
Contact
Questions about scope, a puzzle idea you're unsure about, or getting your club involved as a group — join our Slack community and DM Zac, or email zac@terac.com. We'd rather talk through an idea early than review a finished puzzle that was never going to work.
The Berkeley Stumper Challenge is run by Terac. Terac connects 160,000+ registered experts with research work at the world's most advanced AI companies.