Kaloko

Sinfin Kaloko

Shipped with AI. Now let’s actually look at it.

We all build faster than we review. Screens appear in branches nobody has opened, in languages nobody has clicked through. That is not an AI problem, it is a looking problem. Kaloko walks the acceptance plan, captures every screen and e-mail on the way, checks each criterion and lays the run out as one canvas: the contact sheet of the sprint, with a stamp of who looked at what. For teams that ship with coding agents: developers get evidence, QA and product owners review in one place, clients approve what they saw.

Free plan, free for good · e-mail sign-in · the first run in ten minutes

Here is what we built and what got approved

A live run of Kaloko on its own public pages, refreshed daily. Findings are genuine; when something fails, it is fixed in the next release, not hidden.

Live demo · a real run of Kaloko on itself · 10 of 10 steps passed · 2026-09-29Open the run in Kaloko ↗

Things that happen

  • PR merged at 2:14. Screenshots nowhere.
  • The Slovak version exists. Probably.
  • Who approved this? The bot did.
  • Three agents, four branches, one staging, zero memory of what changed.
  • “Works on my machine” is now “works in my context window.”

How it works

  1. Acceptance plan

    An ACCEPTANCE.md next to the task says what each screen must achieve. A scenario.yml turns it into steps, branches and criteria the CLI can run.

  2. Walk

    A Playwright script, an agent with a browser, or a person walks the flow in a shared browser. Every step is captured the same way: full-page screenshots on desktop and mobile, rendered HTML, console errors, e-mails from the test mailbox.

  3. Check

    Countable criteria (elements, texts, prices, metadata, overflow) are checked in the live page with evidence. Criteria that need judgement get a probability from an AI evaluator that reads a reduced text outline of the page; Kaloko turns it into pass, fail or needs review.

  4. Canvas and stamps

    The run becomes a canvas: screens, branches, locales, evidence. Reviewers approve a step, return it with a note, or decide a manual criterion. The agent reads the notes with kaloko feedback and fixes them in the next run.

Three stamps. The human is the point.

approved

A person looked and agreed.

returned

A person looked and wrote what to change.

needs a human

The evaluator was not sure. Someone decides on the canvas.

What the canvas shows

  • Desktop and mobile screenshots of every step, e-mails and third-party screens
  • Pass, partial or fail per criterion, per locale and viewport, with the evidence
  • Page metadata: title, description, canonical, hreflang, Open Graph, JSON-LD
  • Branch, commit, author, PR and environment; the scenario’s history and a pixel diff against the baseline
  • Replay: step by step with a choice at every decision, for people who do not read canvases
  • Links to a specific step and criterion you can send to colleagues

What teams use it for

Accepting a task

The agent walks the plan on local or review, shares the canvas, posts the link into the PR. The reviewer accepts or returns it there.

Regression before a release

Run the flows against staging, compare with the accepted baseline, see what moved: statuses, criteria, pixels.

Landing pages and SEO on production

Read-only runs of public pages: headings, metadata, hreflang, overflow, console. Kaloko runs this on itself every day.

Agent-driven QA

Claude Code, Codex or CI run the same commands a person would; the canvas is where the two meet.

Fits the tools you already use

  • Claude Code and Codexa skill that knows when and how to walk, capture, evaluate and share
  • GitHubkaloko share --pr keeps one comment per scenario on the pull request
  • CIthe CLI runs headless; a read-only token is enough to read results
  • Mobile and desktop appsAndroid over adb, the iOS Simulator, Electron and macOS apps on the same canvas; screens from Maestro, Appium or XCUITest imported with their UI tree (guide)
  • MCPruns, feedback, verdicts and the inbox over MCP at kaloko.app/mcp
  • E-mail and Slackdigest of what needs you, immediate mail when your run is returned, one Slack channel for the team

For designers: the design development and QA run against

Design the screens of a flow as live HTML with your own agent and skills, together with your team in real time. Accept a version, and every implementation run is checked against it: pixels, structure, tokens and contrast.

Draft, live and shared

Your agent writes each step as live HTML; the team clicks through the prototype, pins comments on elements and sees each other’s cursors.

Versions without duplicates

Publishing makes a version of the whole flow. Unchanged steps are stored and rendered once; compare shows only what moved.

Design reference

An accepted version becomes the reference. Developers’ agents take the live files and tokens from it, step by step.

Checked implementation

The design pack compares every implementation run with the reference and names the token to use where it drifts.

How to design with your agent in Kaloko →

Install once, then work through your agent

Kaloko runs where your code and your agent are. The service stores and versions the results, shows the canvas and collects approvals.

  1. Add the CLI to the project
    npm install --save-dev kaloko

    Needs Node 20 or newer. Update later with npm update kaloko.

  2. Create the config and install the skill
    npx kaloko init --agent claude --org <your-org>

    The skill is copied to .claude/skills/kaloko. npx kaloko doctor checks Chrome, the config and the token.

  3. Create your organization and a token

    Create an organization; you become its admin. The start page offers a tester token in one click, later under Settings → API tokens. Put it into the project .env:

    KALOKO_TOKEN=qwk_…
    TYPESAFE_API_KEY=…   # optional: semantic evaluator

Then just ask your agent

The skill teaches your agent the whole loop: it writes the acceptance plan and the scenario from the task, walks the screens, evaluates, shares the canvas, reads what reviewers said and fixes it. You do not type the commands; you review the canvas.

  • From a task to a shared walkthrough
    Read the task in docs/tasks/TASK-123, write its acceptance plan and a Kaloko scenario, walk it on local and share the result on the pull request.

    Creates ACCEPTANCE.md and qa/scenario.yml, captures every step in each language and screen size, and posts the link.

  • Process the feedback
    Go through the open Kaloko feedback on TASK-123: fix what reviewers rejected, run the affected steps again, share, and answer each comment with what changed.

    Reads threads with kaloko feedback, resolves them with the commit that fixed them.

  • Check a public site
    Draft a Kaloko scenario for https://example.com (home, pricing, contact) with the seo, a11y and perf packs, run it on production read-only and summarise what fails.

    Read-only on production; check packs add SEO, accessibility and performance criteria without writing them by hand.

  • Keep the history tidy
    Clean up old Kaloko runs of TASK-123: show me what a prune keeping the newest 3 would remove, then do it once I agree.

    Always a dry run first; accepted, kept and commented runs stay.

What the agent runs (or run it yourself, e.g. in CI)

The same loop by hand:

npx kaloko start --scenario docs/tasks/TASK-123/qa/scenario.yml --env local
npx kaloko walk        # playwright steps; agent/manual steps: kaloko capture
npx kaloko evaluate
npx kaloko share --pr
npx kaloko feedback    # what reviewers said, with ids to answer

Data & security

The evaluator sees an outline, not the page

The AI evaluator receives a reduced text outline of the page with e-mail addresses and tokens masked. It never receives screenshots or raw HTML.

Sign-in by e-mail link, three roles

People sign in with a one-time link sent to their work e-mail. A viewer reads and reviews, a tester also pushes runs, an admin also manages members, domains and tokens.

Runs expire, files live apart

Every run has an expiry date (30 days by default) and is deleted afterwards. Uploaded files are served from a separate origin under short-lived signed links; captured HTML runs in a sandbox without scripts.

Production stays read-only

A production environment is always read-only in the CLI: public pages as a visitor, no sign-in to back offices, no data created.

Free stays free. Paid plans start with a 14-day trial: card required, cancel before it ends and nothing is charged. Terms of service.

Pricing

Viewers are free and unlimited in every plan: the people who look must never be the bottleneck. You pay for what you ship: shared runs per month, tester and admin seats, retention and storage. Team and Business start with a 14-day trial. For businesses; prices exclude VAT.

Free

Free

free for good

no card needed

For one person and their agent

  • 30 shared runs / month
  • 2 tester + admin seats
  • unlimited viewers
  • 30 days retention · 2 GB storage
  • canvas, replay, compare, inbox
Business

199 € / month

billed monthly

yearly you save 478 €

159 € / month 199 €

billed 1,910 € yearly

you save 478 €

For several teams and documentation

  • 1500 shared runs / month
  • unlimited tester + admin seats
  • unlimited viewers
  • 365 days retention · 100 GB storage
  • accepted and reference runs never expire; documentation layer coming
Enterprise

on request

yearly contract, invoice

SSO, SLA, DPA, self-hosted

Governance and operations

  • unlimited shared runs / month
  • unlimited tester + admin seats
  • unlimited viewers
  • unlimited storage
  • SSO/SCIM, custom retention, own files domain, self-hosted, SLA, DPA, invoice

Semantic evaluation is bring-your-own-key in every plan; local previews are always unlimited. Compare all plans →

Questions we get

Does Kaloko replace tests?

No. Unit and end-to-end tests tell you whether the code works. Kaloko shows what was built and lets a person say whether it is what was meant.

Do I need an AI evaluator?

No. Countable criteria run without one. For judgement criteria bring your own key (the default is JEV by TypeSafe); without it they stay “not evaluated” until a reviewer decides.

What does it cost?

Free stays free. Team is 49 € and Business 199 € per month, −20 % with yearly billing, both with a 14-day trial. Data stays readable and exportable whatever the plan.

Where is the data?

Cloudflare storage chosen by Cloudflare; the operator is a Czech company under EU law. Runs are yours and expire on the date you set.

What comes next

A weekly “what changed” canvas assembled from runs, check packs for SEO and accessibility, and a GitHub check that turns green when the run is accepted. We tell admins before anything changes that affects them.