Assessments for the AI era

See how engineers work with AI.

Send a real coding task. See how candidates work with their AI. Get scores that cite the evidence.

0evals per submission0private repo per candidate0answer board per candidate0keystrokes logged
  1. 1Candidate works
  2. 2Evals score
  3. 3You hire
Claude Code

Rejected the agent's edit

Screen recording

42:18 of real work

Code walkthrough

Explained the paging fix

Telos evals–
  • Code
  • AI collaboration
  • Process
PPriyaBackend Engineer
Strong hireOnsite

Traditional interviews stopped measuring anything.

When every candidate's agent can ace the test, a perfect score tells you nothing. Telos measures how they actually work.

LeetCode & take-homes

On traditional interviews nearly every candidate scores 100.

Everyone scores 100.

  • Gameable: an agent solves it in seconds
  • Cheatable: no one sees how the answer arrived
  • Not the job: puzzles, not shipping software
Telos

On Telos the same candidates spread across the whole range.

Now you can tell them apart.

  • A real task in a real repo
  • Every prompt, test and edit captured
  • Every score cites the moment behind it
How it works

From invite to decision in four steps.

1

Send a real task.

Write the brief and invite by email. Each candidate gets a private repo.

Task
acme/assess-priyacommits
2

They work with AI, on the record.

Their own editor and Claude Code. Timed runs record screen and mic.

Your editorClaude Code
3

We process everything.

Commits, the screen recording and every agent prompt, read in order.

GitHub commits14 commits
Screen recording42:18 with mic
Agent prompts38 prompts · 1.2M tokens
4

Get scores with receipts.

The repo locks, the evals run, and every score cites its evidence.

Code
AI collaboration
Process

Their tools, captured.

Candidates work in the editor and agent they already use. Capture ships inside the assessment repo, and everything lands on one answer board.

orders.ts — assess-priya
Explorersrc/orders.tsorders.test.tscursor.ts
orders.tsorders.test.ts
12export async function listOrders(cursor?: string) {
13 const page = await db.orders.after(cursor).take(50);
14 return { items: page, next: page.at(-1)?.id };
15}
Terminal$ bun test✓ 14 passed
mainTelos · reporting

Cursor or VS Code

The Telos extension reports what they do by hand: saves, terminal commands, and how they review each diff. Keystrokes never leave the editor.

savesterminal commandsdiff review
assess-priya — claude

Claude Code in the terminal

Capture settings ship inside the repo. Any Claude Code session started there reports automatically, including every edit the agent proposed and whether they kept it.

committed hooksper-edit diffssecrets redacted

Company-paid AI keys. Give each candidate one key with a budget. It stops at the deadline, on submit, or once the budget is spent.

Assessment types

Pick the format that fits the role.

Each format is set per assessment, and candidates see what's on before they accept.

Mon
Tue
Wed
Thu
Fri
Due Fri 17:00

Take-home

Their own pace, up to your deadline. They work in a private repo and submit when it's ready.

own pacedeadlineprivate repo
1:24left
REC

Timed

The clock starts when they press Start, not when they accept. Screen and mic are recorded throughout, and the work is collected the moment time's up.

starts on Startauto-collectedalways recorded
DFrustrated customerDouble charge02:14

Job simulation

Next up

Put them on a live call with an AI customer: a double charge, a late delivery, an outage. See how they handle it.

AI voice callerscenario + persona
Evaluation

Every score cites its evidence.

Three evals run on every submission. Each claim links to the lines, prompt or commit it came from. A score with no evidence doesn't ship.

41export async function listOrders(cursor?: string) {
42 if (cursor && !isCursor(cursor)) throw new BadCursor();
43 const page = await db.orders.after(cursor).take(50);
44 return { items: page, next: page.at(-1)?.id };
45}
cited · orders.ts:42–44 · error handling

Code

How good is the code they wrote? An overall judgement across structure, readability, error handling and tests, reviewed by an agent that reads the repo.

1
4
6
4 of 7 prompts checked before moving on

AI collaboration

How well did they work with AI? Context given, how they split the task, how they checked the output, how they recovered.

commits per hour · small, steady steps

Engineering process

How did they run the work? Small steps, commit hygiene, priorities.

Evals
What was delivered
Working with AI
Recording & follow-up

Answer board

One page per candidate: evals, what they delivered, how they worked with AI, recording and follow-up.

Code
AI
Process

Insights

Compare everyone who took the same assessment: spreads, not anecdotes.

engineering-handbook.pdf
docs.acme.dev/style-guide
review-checklist.md
How do we review pull requests?

Your company context

Upload docs or add URLs. The copilot searches them, so its answers reflect how your team actually works.

Copilot

Ask your hiring anything.

One copilot across your whole workspace. It reads your roles, candidates, sessions and company context, and cites every fact it states. When you ask it to act, it proposes the change and waits for you.

  • @mention candidates, roles and assessments
  • Nothing changes until you confirm
  • Read-only MCP access from Claude Code

Priya ran the tests before 5 of 6 commits and rejected 3 agent editsprompt #7orders.ts:42. Sam committed agent output without running it twicecommit a1f3c09.

Proposed actionSet Priya → Shortlisted?
Trust

Transparent by design.

Before they start, candidates read exactly what is captured and what isn't. Nothing is recorded until they accept.

What you see

  • Prompts, agent replies and the tools the agent ran
  • Every edit the agent proposed, and whether they kept it
  • Terminal command lines, never their output, with secrets removed
  • Commits and pushes to the assessment repo
  • Screen and microphone from Start to submit in timed assessments

What you never see

  • Keystrokes, or anything outside the assessment repo
  • The model's private reasoning
  • Their Claude account, billing or credentials
  • Anything after they submit
disclosure accepted before startingversioned, so the terms never shiftsecrets redactedrecordings expire automatically

Got an assessment invite?

Sign in with the email it was sent to. Work in your own editor and Claude Code, and see exactly what's shared before you start.

Open my assessments

Hire on evidence.

Send your first assessment today and see how your candidates actually work.