How to Compare AI Tools for One Real Task: A Student Evaluation Project
Compare AI tools by fixing one real task, inputs, constraints, rubric, and test cases before opening any tool. Run the same cases, preserve outputs, verify factual claims, record.

Compare AI tools by fixing one real task, inputs, constraints, rubric, and test cases before opening any tool. Run the same cases, preserve outputs, verify factual claims, record model and date, review privacy and cost conditions, and make a recommendation only for that task.
The project produces a reproducible test set, scoring sheet, evidence log, limitation register, and decision memo. It replaces popularity claims with observable strengths and failures.
Do not upload confidential data, claim a universal best tool, or treat one run as a stable benchmark. Provider features, models, limits, and prices can change, so date the evaluation.
What You Need Before You Start
Choose a permission-safe task with an answer you can verify, such as summarizing and extracting actions from three public policy documents.
- Access to selected tools under their current terms
- A fixed public test set
- A spreadsheet and secure output folder
- An independent method for checking facts
Keep these boundaries in place:
- No sensitive data
- No universal leaderboard
- No invented cost or model capability
1. Bound the task and success conditions
Action and reason: Define user, input, required output, prohibited behavior, time limit, and five test cases. A broad productivity comparison cannot produce a defensible decision.
Inputs: The public documents and a reference answer. Treat these inputs as the working boundary; adding an unreviewed dependency changes the task and should trigger a new check.
Expected output: A frozen evaluation protocol. Keep the artifact with the project so another person can inspect the decision instead of relying on a finished screenshot.
Verification: Ensuring every criterion can be observed or checked. Record the input, expected result, observed result, and decision. A pass is credible only when another person can repeat the same check.
Failure condition: The task changes after seeing a favored tool’s output. Correction: Freeze protocol and rerun all tools after any justified change. Repeat the original verification after the correction and retain the failed observation as part of the evidence trail.
2. Build a weighted rubric
Action and reason: Assign accuracy, completeness, instruction following, traceability, usability, latency, privacy, and cost roles. Fluent writing must not outweigh factual errors.
Inputs: Task risks and user priorities. Treat these inputs as the working boundary; adding an unreviewed dependency changes the task and should trigger a new check.
Expected output: Scoring definitions with examples of pass, partial, and fail. Keep the artifact with the project so another person can inspect the decision instead of relying on a finished screenshot.
Verification: Having a second reviewer score one sample and reconcile differences. Record the input, expected result, observed result, and decision. A pass is credible only when another person can repeat the same check.
Failure condition: Criteria overlap or subjective style dominates. Correction: Rewrite definitions and reduce weight on nonessential preference. Repeat the original verification after the correction and retain the failed observation as part of the evidence trail.
3. Control prompts and settings
Action and reason: Use equivalent instructions, attachments, model settings, and fresh sessions. Changing conditions invalidates comparison.
Inputs: The frozen prompt, file set, and run log. Treat these inputs as the working boundary; adding an unreviewed dependency changes the task and should trigger a new check.
Expected output: Dated runs with tool, model, settings, input, and raw output. Keep the artifact with the project so another person can inspect the decision instead of relying on a finished screenshot.
Verification: Checking logs before scoring. Record the input, expected result, observed result, and decision. A pass is credible only when another person can repeat the same check.
Failure condition: One tool receives clarifying help that others did not. Correction: Create a documented second-round protocol and apply it equally. Repeat the original verification after the correction and retain the failed observation as part of the evidence trail.
4. Verify outputs and repeat unstable cases
Action and reason: Check extracted facts against source passages and rerun cases with material variance. Plausibility is not evidence of accuracy.
Inputs: Source documents, reference answer, and claim log. Treat these inputs as the working boundary; adding an unreviewed dependency changes the task and should trigger a new check.
Expected output: Verified, unsupported, contradicted, and omitted claim labels. Keep the artifact with the project so another person can inspect the decision instead of relying on a finished screenshot.
Verification: Linking every scored factual issue to source evidence. Record the input, expected result, observed result, and decision. A pass is credible only when another person can repeat the same check.
Failure condition: Reviewers score from memory or accept citations without opening them. Correction: Perform source-level checks and lower confidence where verification is unavailable. Repeat the original verification after the correction and retain the failed observation as part of the evidence trail.
5. Review privacy, limits, and cost context
Action and reason: Read current provider documentation for data handling, file limits, account terms, and pricing basis. A useful output may still be unsuitable for the user’s data or budget.
Inputs: Official provider pages accessed on the evaluation date. Treat these inputs as the working boundary; adding an unreviewed dependency changes the task and should trigger a new check.
Expected output: A constraints register separate from quality scores. Keep the artifact with the project so another person can inspect the decision instead of relying on a finished screenshot.
Verification: Citing each current condition and marking unknowns. Record the input, expected result, observed result, and decision. A pass is credible only when another person can repeat the same check.
Failure condition: Assumptions from another plan or old model are reused. Correction: Recheck official documentation and narrow the recommendation. Repeat the original verification after the correction and retain the failed observation as part of the evidence trail.
6. Write a task-specific decision memo
Action and reason: Summarize results, failure patterns, tradeoffs, uncertainty, and the chosen tool for this task. A responsible comparison ends with a bounded decision, not a universal ranking.
Inputs: Scores, evidence, constraints, and repeat results. Treat these inputs as the working boundary; adding an unreviewed dependency changes the task and should trigger a new check.
Expected output: A recommendation with conditions and fallback. Keep the artifact with the project so another person can inspect the decision instead of relying on a finished screenshot.
Verification: Confirming that each conclusion traces to recorded evidence. Record the input, expected result, observed result, and decision. A pass is credible only when another person can repeat the same check.
Failure condition: The conclusion generalizes beyond the tested documents and settings. Correction: Rewrite scope and state what a future evaluation must test. Repeat the original verification after the correction and retain the failed observation as part of the evidence trail.
Review the Evidence Before You Call It Complete
Run the work as a review, not as a presentation. Start with the promised outcome: Protocol and rubric were fixed before testing. Ask a second person to follow the documented inputs and checks without receiving a private explanation. Record where they cannot reproduce a result, where a decision lacks evidence, and where the artifact depends on hidden knowledge. Those gaps are part of the work and should be corrected before screenshots or portfolio copy are finalized.
Use the remaining acceptance criteria as release conditions: Runs preserve model, date, settings, and outputs; Factual scores link to source evidence; Privacy and cost conditions cite official pages; Recommendation applies only to the tested task. A failed condition should identify the smallest upstream step that owns the defect. Correct that step, repeat the same check, and preserve the before-and-after result. This review discipline is what turns an exercise into credible evidence of skill without claiming client experience, production success, or testing that did not occur.
Completion Standard
The evaluation passes when another student can repeat it and understand why the recommendation is limited.
- Protocol and rubric were fixed before testing
- Runs preserve model, date, settings, and outputs
- Factual scores link to source evidence
- Privacy and cost conditions cite official pages
- Recommendation applies only to the tested task
Students in Gujrat can use Pakistan-relevant public documents and English or Urdu material where the chosen tools support them, but language and regional results must be scored separately rather than assumed from English performance.
When you want guided review of the complete workflow, the Artificial Intelligence course provides a structured path from fundamentals to supervised project evidence. The article remains a self-contained method; the program is the next step for learners who need feedback, correction, and repeated practice.
Sources And Verification Notes
- NIST AI Risk Management Framework: Supports measurement, documentation, limitations, and risk-aware evaluation.
- OpenAI models documentation: Provides current first-party model and capability context for any OpenAI tool included.
- Responsible use of GitHub Copilot Chat: Supports human review, testing, and caution with generated outputs.
FAQ
How many AI tools should a student compare?
Three is usually enough to practise a controlled comparison without making evidence unmanageable. Depth matters more than a long list.
Should price be part of the score?
Record cost as a constraint tied to the tested usage and current plan. Do not combine it blindly with accuracy when the user’s risk makes accuracy non-negotiable.
Can one prompt fairly test every tool?
Use equivalent task requirements and document unavoidable interface differences. If a second-round clarification is allowed, apply the same rule to every tool.
Want to Build Practical Technology Skills?
Explore RisingEdge courses designed to help students learn real skills, build projects, and prepare for career opportunities.



