Assessments for the AI era

See how engineers work with AI.

Every candidate uses an agent now. Kismet sends a real assessment in a private repo, captures how they drive the agent (prompts, edits, commits), and returns evals where every score cites the lines it rests on.

0evals per submission0private repo per candidate0answer board per candidate0keystrokes logged
assess-priya · Backend Engineercandidate's own Claude Code
Session
  • prompt“Add pagination to /orders, keep the cursor API”00:04
  • agentedited 3 files · +84 −1200:05
  • yourejected edit in orders.ts00:07
  • terminalbun test ✓ 14 passed00:09
  • commitfeat: cursor pagination00:11
Evalsevery score cited
Code qualityLevel 0/5
orders.ts:42–61
AI collaborationLevel 0/5
rejected edit · 00:07
Engineering processLevel 0/5
commit a1f3c09 · 00:11
consented, versioned disclosurerepo locked at submit

The code isn't the signal anymore. The process is.

Any candidate's agent can produce working code. You're hiring for how they got there: what they asked for, what they checked, and what they changed.

What a take-home tells you todayWhat Kismet shows you
“Tests pass”
bun test run 6× · first green at 00:41
“Clean, well-structured code”
3 agent edits rejected · 2 rewritten by hand
“Solved it in 3 hours”
14 prompts · 4 commits · steady increments
“Used AI responsibly”
Checked the agent's output before every commit
How it works

From role to decision in four steps.

1

Create a role and an assessment.

Write the brief, then choose take-home or timed, where they work, what gets scored, and whether there's a follow-up.

Name
Format
Timing
Environment
Evals
AI keys
Follow-up
2

Invite by email.

Each candidate gets a private repo made from your template through your GitHub App.

Assessment · Backend Engineer
acme/assess-priya
3

They work the way they work.

In their own editor and Claude Code. Capture ships inside the repo.

Cursor · VS Codeclaude
4

Review the evidence.

At submit or deadline the repo locks and the evals run. One answer board per candidate, Insights across the cohort.

Code quality
AI collaboration
Process

Their tools, captured.

Candidates work in the editor and agent they already use. Capture ships inside the assessment repo, and everything lands on one answer board.

orders.ts — assess-priya
Explorersrc/orders.tsorders.test.tscursor.ts
orders.tsorders.test.ts
12export async function listOrders(cursor?: string) {
13 const page = await db.orders.after(cursor).take(50);
14 return { items: page, next: page.at(-1)?.id };
15}
Terminal$ bun test✓ 14 passed
mainKismet · reporting

Cursor or VS Code

The Kismet extension reports what they do by hand: saves, terminal commands, and how they review each diff. Keystrokes never leave the editor.

savesterminal commandsdiff review

Claude Code in the terminal

Capture settings ship inside the repo. Any Claude Code session started there reports automatically, including every edit the agent proposed and whether they kept it.

committed hooksper-edit diffssecrets redacted

Company-paid AI keys. Give each candidate one key with a budget. It stops at the deadline, on submit, or once the budget is spent.

Timed & recorded

Timed when it matters. Recorded when you need it.

Each is a per-assessment setting, and candidates see which ones are on before they accept.

1:24left

Timed assessments

The clock starts when they press Start, not when they accept. When time's up, the work is collected automatically.

Recording screen and mic42:10
promptcommit

Screen recording

Whole screen and microphone from Start to submit, with a pre-flight check and alerts if the recording drops.

Q1Q2Q3

Recorded follow-up

After they submit, they reopen their own code and walk you through it, one question at a time.

Evaluation

Every score cites its evidence.

Three evals run on every submission. Each claim links to the lines, prompt or commit it came from. A score with no evidence doesn't ship.

41export async function listOrders(cursor?: string) {
42 if (cursor && !isCursor(cursor)) throw new BadCursor();
43 const page = await db.orders.after(cursor).take(50);
44 return { items: page, next: page.at(-1)?.id };
45}
cited · orders.ts:42–44 · error handling

Code quality

Is the code they wrote well made? Structure, readability, error handling and tests, reviewed by an agent that reads the repo.

1
4
6
4 of 7 prompts checked before moving on

AI collaboration

How well did they work with AI? Context given, how they split the task, how they checked the output, how they recovered.

commits per hour · small, steady steps

Engineering process

How did they run the work? Small steps, commit hygiene, priorities.

Evals
What was delivered
Working with AI
Recording & follow-up

Answer board

One page per candidate: evals, what they delivered, how they worked with AI, recording and follow-up.

Code
AI
Process

Insights

Compare everyone who took the same assessment: spreads, not anecdotes.

engineering-handbook.pdf
docs.acme.dev/style-guide
review-checklist.md
How do we review pull requests?

Your company context

Upload docs or add URLs. The copilot searches them, so its answers reflect how your team actually works.

Copilot

Ask your hiring anything.

One copilot across your whole workspace. It reads your roles, candidates, sessions and company context, and cites every fact it states. When you ask it to act, it proposes the change and waits for you.

  • @mention candidates, roles and assessments
  • Nothing changes until you confirm
  • Read-only MCP access from Claude Code

Priya ran the tests before 5 of 6 commits and rejected 3 agent editsprompt #7orders.ts:42. Sam committed agent output without running it twicecommit a1f3c09.

Proposed actionSet Priya → Shortlisted?
Trust

Transparent by design.

Before they start, candidates read exactly what is captured and what isn't. Nothing is recorded until they accept.

What you see

  • Prompts, agent replies and the tools the agent ran
  • Every edit the agent proposed, and whether they kept it
  • Terminal command lines, never their output, with secrets removed
  • Commits and pushes to the assessment repo
  • Screen and microphone from Start to submit, only if the assessment records

What you never see

  • Keystrokes, or anything outside the assessment repo
  • The model's private reasoning
  • Their Claude account, billing or credentials
  • Anything after they submit
disclosure accepted before startingversioned, so the terms never shiftsecrets redactedrecordings expire automatically

Got an assessment invite?

Sign in with the email it was sent to. Work in your own editor and Claude Code, and see exactly what's shared before you start.

Open my assessments

Hire on evidence.

Send your first assessment today and see how your candidates actually work.