Ember

The Ember Codex

The box system: one prompt and one eval per section of your report

Every box in an Ember template holds one prompt and one eval, and a box that misses the standard you set is reworked before the report reaches you.

Handy Gunawan · · about 3 min to read

Ember Console is a session-report builder for coaches, in closed beta at $8 a month with 5 free sessions a month. Every box in its template holds one prompt and one eval. A box that misses the standard you set gets reworked before the report reaches you, so the checking happens during the build.

Paste a transcript into a chat window instead, and the only check is you, reading the finished file.

How does Ember build a report, step by step?

Your AI does the writing, and Ember supplies the workflow around it.

  1. Ember records the session, or takes a recording you have.
  2. You reach Ember from ChatGPT, Claude or Littlebird.
  3. Ember hands over the box prompts and the eval system.
  4. Your AI fills each box of your template.
  5. A box that fails its eval is reworked until it passes.
  6. You edit the finished report in the Ember dashboard.

Your letterhead rides on the template, and your edits feed the next report. Ember’s landing page puts that part this way: “Get tailored AI output for every section of your template, while Ember learns from your edits and checks each result against your own criteria.”

Why does a prompt alone leave the checking to you?

Because a prompt has nowhere to keep your standard, so nothing compares a section to it until you do.

An instruction is the part of the job you can hand someone, so a search for session-note prompts returns prompts. Coach Guido gives away one that turns notes into a one-page prep brief, and Amazon sells a pack of 59 done-for-you prompts for therapists.

Checking one model’s output against written criteria has a name outside coaching software: Langfuse’s documentation calls it LLM-as-a-judge. I have not published what Ember’s eval inspects, so read a box as the place your standard lives and nothing more.

A check reports only on what it measures, which is how a symptom score misses a better week.

What is not stated yet?

Four questions I have not answered anywhere public, and I will not guess at them here.

  1. What a box is, exactly.
  2. How many boxes a default template has, and whether you can add, remove or reorder them.
  3. What the eval checks: length, wording, the presence of a fact, or a rule you typed.
  4. What happens when a box fails: how many reworks, and what you see while it runs.

Until those answers exist, the box system is a mechanism to inspect rather than a claim to take from me.

What you can check this week is narrower. Type your standard into a box, read the section that comes back, and see whether they match. That is the test I would want run on my own software.

Questions people ask next?

Can ChatGPT write a coaching session report from a transcript?

Yes, and you stay the only check. A single prompt gives one draft, and nothing in the run holds a section against the standard you keep. Ember sends a prompt and an eval per box, then reworks the box that misses.

How do you make AI follow a report template?

Split the template and give every section its own instruction and its own check. One r/ChatGPT thread describes ChatGPT mangling simple templates (r/ChatGPT). What my eval inspects is still unpublished.

What does Ember Console cost in the closed beta?

$8 a month, with 5 free sessions a month. It is a session-report builder for coaches, and the beta is closed, so access is not open yet. The landing page names ChatGPT, Claude and Littlebird.

Ember Console is a session-report builder for coaches, in closed beta at $8 a month with 5 free sessions a month. What it does is at https://ember.anyeads.com/

Start free →