Build standard

How a skill is built. And why it works.

A skill is your business knowledge written down so an AI agent follows it the same way every time. Most are a single page of instructions. Every WalterSignal skill uses the same layout, and each part of it exists to stop a specific way agents go wrong.

01 · Always in view

Every header

  • invoice-check
  • quote-buildermatch
  • crew-schedule
  • intake-triage
  • inventory-count

A few lines per skill. The agent reads these for every request and picks the one whose description fits.

02 · Opened on a match

SKILL.md

  • Procedure and gates
  • Rules, each with its reason
  • Checklist that defines done

The full instructions load only for the skill that matched.

03 · Opened at the step

references/ · scripts/

  • pricing-policy.mdreference
  • materials.mdreference
  • calc_quote.pyruns as code
  • check_totals.pyruns as code

Detail is read when a step needs it. Exact work runs as code.

04 · Before you accept it

evals/

  • Fixed promptsyour real work
  • Pass/fail checklistwritten first
  • Run without the skill
  • Run with the skill

Both runs are graded on the same checklist, and you see the difference.

The agent carries only what the current step needs, so the instructions that matter are never crowded out.
The header

The part the agent always reads.

An agent with many skills cannot read all of them for every request. It keeps only each skill’s header in view and opens the full file when a header matches the job. The header decides whether the skill runs at all, so it is managed field by field.

---
name: quote-builder
description: Build a customer quote from a site visit. Use when someone
  asks to "quote this job", "price the visit", or pastes visit notes.
  Not for changing a quote that was already sent.
version: 1.2.0
argument-hint: "[visit notes or job id]"
---
name
A short, fixed identifier, such as quote-builder.
Why: People and other skills call the skill by this name, so it never changes once in use.
description
One or two sentences on what the skill does, the exact phrases that should trigger it, and the jobs it is not for.
Why: The agent keeps every skill's header in view at all times and loads the rest of the file only when it picks that skill. The description is the only thing it reads when choosing, so it is written to be precise and kept short.
version
A version number raised on every change to the skill.
Why: Test results are recorded against a version, so you always know which version passed and what changed since.
argument-hint
For skills a person runs directly, the input the skill expects.
Why: Anyone on your team can run the skill as a command without reading its instructions first.
01

A header that decides when it runs

Every skill opens with a short header: its name, a description of when to use it and when not to, and its version. The header is set out in full above.

Why it works

The agent chooses a skill by reading that header. A precise one means the skill is used every time the job comes up and stays out of the way when it does not.

02

A short procedure, in order, with gates

The main file sets out the work as numbered phases. A gate between phases says what must be true before the next one starts, such as confirming the brief before any building begins.

Why it works

Agents skip steps when the order is implied. Written gates make the order explicit, so the expensive work never starts on a wrong assumption.

03

Rules that carry their reason

Each hard rule states the failure it prevents, often the real incident that produced it.

Why it works

An agent that knows why a rule exists applies it correctly to cases the rule did not foresee. You can also read the reason and judge whether the rule still fits your business.

04

Detail kept in reference files

Long material such as product specs, policies, worked examples, and style guides sits in separate reference files. The main file says which one to open at which step.

Why it works

An agent works best with only what the current step needs in front of it. Detail is there when needed without crowding out the instructions that matter.

05

Scripts for anything that must be exact

Calculations, file conversions, data checks, and quality gates are written as scripts the skill runs, often one command that runs every gate.

Why it works

Code returns the same answer every time; a model asked to do arithmetic or remember ten checks does not. The agent's judgment is kept for the parts that need judgment.

06

A checklist that defines done

Every skill ends with pass/fail checks the agent must meet before it reports the job finished, plus a list of the mistakes it is known to make.

Why it works

"Looks good" is not a standard. A checklist turns done into something that can be checked, by the agent and by you.

07

Tested before you accept it

A fixed set of your real prompts and a pass/fail checklist are written before the skill. The same agent runs every prompt with and without the skill, and the scores are compared.

Why it works

You see a measured difference instead of a demo. If the skill misses an item, the skill is fixed and re-run; the test never changes to fit the result.

08

Headers tuned by measurement

The description in each header is tested against requests that should and should not trigger the skill, and rewritten until it picks correctly.

Why it works

A skill that does not trigger is worth nothing, and one that triggers on the wrong job does harm. Triggering is measured, not guessed.

09

Small skills that work together

A large job is an entry skill that calls smaller specialist skills in sequence, each with its own checks.

Why it works

Each skill stays small enough to test on its own, and improving one does not risk breaking the others.

Self-improving

Every miss makes the skill better.

A skill is not finished at handover. Each mistake it makes is turned into a rule and a test, so it is fixed once and checked forever. The agent proposes the fix; a person approves it before it ships.

  1. Step 1

    Catch the miss

    Any time the skill gets something wrong, in testing or in daily use, the case is recorded with what the right answer was.

    Why: A mistake that is only fixed in the moment comes back next week. A recorded one can be fixed for good.

  2. Step 2

    Write it into the skill

    The fix goes into the skill as a rule that states the mistake it prevents, or into a script when the step must be exact.

    Why: The next run, and every agent that uses the skill, inherits the lesson instead of relearning it.

  3. Step 3

    Add it to the tests

    The case that failed becomes a new test prompt with its own pass/fail items. Earlier prompts are never edited, so scores stay comparable over time.

    Why: The same mistake cannot return without a test failing first.

  4. Step 4

    Re-run everything, then release

    The whole test set runs again. The new version is released only if it fixes the miss and nothing that passed before has broken.

    Why: Improving one answer never quietly breaks another, and every version has a score you can see.

Common questions

The system around the skills.

Who owns the skills and the system?

You do. Skills are plain files you can read and edit, and the code, the domain, and the data are yours outright, with no per-seat license. Hosting can sit in your own accounts from the start, or WalterSignal can host and operate it for you and move it into your accounts on request.

Is my data separated from other clients?

Yes. Every client runs as their own deployment with their own database. Two clients never share a database, so a fault in one system cannot reach another.

Is my system a one-off build?

No. Every client runs on the same maintained codebase, switched on through a configuration file that selects modules such as ordering, scheduling, CRM, quoting, and billing. Fixes and improvements reach every client without a rebuild.

Where is security enforced?

In the database. Every table denies access unless a rule allows it, public forms write only through a function that re-checks every field, and prices are calculated by the database rather than taken from the browser. A test suite checks every table for these rules.

How are changes released?

Every database change is a versioned file, and every code change runs type checks, tests, and structural checks. Releases go through one deploy command that checks the target and settings and refuses to start if anything is wrong.

Which AI model does it use?

Whichever fits. AI features call one standard endpoint set by configuration, so changing the model or provider, or moving to your own hardware, is a settings change. The AI proposes and a person confirms before anything is saved.

Can another developer maintain it?

Yes. It is built on mainstream tools: Next.js, Postgres through Supabase, and Vercel. The handover includes the full source and documentation.

What can an AI agent do to my data?

Read it, by default. Agent connections to your database open in read-only mode, so the database itself refuses any write. Write access is granted by a person for a specific job.

Can anything go out without my approval?

No. Every email, publication, or code merge an agent prepares is shown to you as a proposal first, and it runs only when you approve it.

Where are passwords and API keys kept?

In a password vault, read only when the system starts. They never appear in source code, skill files, or configuration, and every change is checked for credentials before it is committed.

See it running

Ask to see a skill and its test results.

On a call, a working skill can be opened file by file and its with-and-without results shown. If it cannot be shown, it does not belong here.