CUA Evaluation Contributor — Job Description & Engagement Terms
1. Position Overview
The CUA Evaluation Contributor is engaged to evaluate model-generated computer-use (CUA) trajectories on the OpenCUA project using the SuperAnnotate platform, in accordance with the project SOP and rubric. A Contributor may be assigned to one or more of three roles — Base Annotator, QC Annotator (Reviewer), or Audit Annotator — as directed by the Team Lead. All roles apply the same evaluation rubric; they differ in stage, scope, and independence.
2. Project Context
Each task consists of a goal, a subgoal list with a corresponding App Used list, and a multi-step trajectory (screenshot, action, action JSON, and reasoning per step). The Contributor evaluates how well the model understands the goal, decomposes and progresses through subgoals, reasons coherently, acts appropriately, and maintains safety — judging every element against the visible UI state and the project rubric. All work is performed inside SuperAnnotate; all judgments and feedback are entered directly into the tool.
3. Common Responsibilities (All Roles)
Regardless of assigned role, the Contributor shall:
- Apply the project rubric objectively, grounding every judgment in the screenshot, reasoning, executed action, and action JSON.
- Complete all mandatory fields for the assigned role and follow the prescribed callout and justification structure.
- Trigger Verify Submission and resolve all validation errors and missing mandatory fields before submitting any task.
- Maintain the independence of blind evaluations where the workflow requires it, and never modify a submitted blind evaluation.
- Meet the quality, consistency, and turnaround standards defined by the SOP and the Team Lead.
4. Role-Specific Scope of Work
4.1 Base Annotator (Primary Evaluation)
Performs the first full independent evaluation of every assigned task.
- Evaluate Goal Clarity, Specificity & Safety (1–5); skip with the appropriate reason if the score is 4 or below.
- Validate and correct the Subgoal List and App Used list so they are MECE, sequentially dependent, and positionally aligned (one App Used entry per subgoal).
- For every step, assign Reasoning Coherence (1–5) and evaluate Subgoal Targeted, Subgoal Feasibility, Subgoal Progression, and Step-Level Subgoal Completion (feasibility, progression, and completion judged only for the latest targeted subgoal).
- Assign final Subgoal Fulfilment (Fully / Partially / Not Completed) per subgoal, documenting all attempts, step ranges, outcomes, and reasons.
- Evaluate Trajectory Safety / Harmfulness (1–5).
- Provide callouts for all scores of 3 or below and justifications for all scores of 4.
4.2 QC Annotator / Reviewer (Second, Blind Evaluation)
Independently re-evaluates 100% of tasks and adjudicates against the Base evaluation.
- Complete a full blind step-level evaluation using the same subgoal list and rubric, without visibility into the Base Annotator's work; submit to freeze it.
- After the Base evaluation is revealed, compare both across all step-level and task-level parameters and record an Agree/Disagree decision for each applicable parameter.
- Provide clear, evidence-based feedback for every disagreement.
- Apply the rework decision: return the task if step-level disagreement exceeds 5%, or if there is any disagreement on a task-level parameter (Goal Clarity; Subgoal List Correctness; Subgoal Fulfilment; Trajectory Safety); otherwise pass it forward.
4.3 Audit Annotator (Final Validation, Sampled)
Runs a final independent blind pass on a sampled subset (~20%) of completed tasks to measure reviewer agreement; does not trigger rework.
- Independently evaluate every step in the sampled task, limited to the five step-level parameters (Reasoning Coherence, Subgoal Targeted, Feasibility, Progression, Step-Level Completion), using the completed corrected subgoal list; submit to lock.
- Compare against the Base evaluation and review QC's visible Agree/Disagree decisions for context.
- Compile a single consolidated task-level feedback summary covering only confirmed disagreements; enter it in the tool and paste it identically into the offline daily working sheet.
- Select the Audit decision from the tool-calculated disagreement percentage (Agree/Thumbs Up if ≤5%; Disagree/Thumbs Down if >5%) and submit via Audit_Complete regardless of outcome.
- Out of scope: Goal Clarity, Subgoal List/App Used, Sequential Dependency, Subgoal Fulfilment, Trajectory Safety, per-step written feedback, and any rework loop.
5. Quality Standards
- Two trained contributors applying this SOP and rubric to the same task should reach materially the same conclusion. Scores must reflect rubric definitions, not personal preference or style.
- Every low score (≤3) requires a callout; every score of 4 requires a justification; all comments must be specific, evidence-based, tied to the rubric, and written in the prescribed structure.
- A task is complete only when all applicable fields are filled, all rationales documented, structural correctness confirmed, and Verify Submission has passed.
6. Compensation
Work is compensated on a pay-per-task basis, determined by the task's trajectory-length bucket (step count) and the Contributor's assigned role (Base Annotator, QC Annotator, or Audit Annotator). Each task is paid at the rate specified for its bucket and role in the attached Rate Schedule, under the applicable Tech or Non-Tech track as classified by the Team Lead.
Buckets range from 1–25 steps through 226–260 steps, with the per-task rate increasing as trajectory length increases. The complete bucket-wise rates are set out in the attached Rate Schedule, which forms part of this agreement.
Approved tasks only. Only approved tasks are counted for payment. A task is treated as approved once it has passed all required stages — through Review, and through Audit for the 20% sampled tasks — and meets the client quality standards shared with contributors. Tasks that have not cleared these stages and standards are not payable.
One payment per role, per task. The per-task rate already incorporates the full Average Handle Time (AHT) for the role, including all stages, rework, and alignment loops. Any effort spent on reworking or aligning a task is treated as part of that task's AHT and is not separately compensated. Accordingly, each role is paid once per task — a Base Annotator is paid once for annotating a task, a QC Annotator once for reviewing it, and an Audit Annotator once for auditing it (where the task falls within the audit sample).
Detailed bucket and role wise rates are attached here in this sheet and in Appendix I.
https://docs.google.com/spreadsheets/d/1cEXGBOAK7cu3gws8BGq2aUdqNRunTh4bzfwf_QPErfY/edit?usp=sharing
7. Tooling
- Annotation platform: SuperAnnotate.
- The Contributor shall use only the approved tools and record time and task completion as directed.