The Ember Codex
Setting the minimum standard: how to write an Ember Console eval a box can fail
An eval is the standard a box has to clear before you read the report. Three tests make one a reader can mark pass or fail.
In Ember Console, the session-report builder for coaches I build, each box of your report has a prompt and an eval. The eval is the standard that section clears before you read it. Ember Console is in closed beta.
“Make it warm” is the eval a coach types first. It says exactly what you want, and no box can fail it.
Three tests turn that into a standard. Writing the box’s prompt is a separate job, covered in the GROW questions build.
What makes an Ember Console eval checkable?
A reader who was not in the session can mark it pass or fail.
- Can a reader tell pass from fail?
- Does it name this box and nothing else?
- Does it name the weak note you never want again?
“Make it warm” fails the first test and the third, because warmth is in the reader. “No commitment appears without the date the client gave it” passes all three, and a coach reading the box can point at the miss.
Nielsen Norman Group made the harder point in August 2026, in a page titled “One AI Output Is an Example, Not an Evaluation”: ask how often the output is good, not whether this one is.
How do I name the weak note I never want to see again?
Take the last report that made you wince, find the sentence, and write the rule it breaks.
Coach Alicja Nocon, who coaches neurodivergent employees, makes the timing argument: “receiving feedback twice a year on something that happened up to six months ago simply does not work.”
A standard you apply only when you reread last quarter’s reports has the same problem, and the one inside Ember Console runs on every build. The landing page describes it this way: “Get tailored AI output for every section of your template, while Ember learns from your edits and checks each result against your own criteria.”
Every time an AI fills and builds the report it receives those prompts and the eval system, and reworks each box until it passes the minimum standard you set, which is the box system. A test behind every section is also how a life coaching template comes together.
What is not stated yet?
Four questions about a box are still open in Ember Console, and the answers are mine to publish.
- What a box is: a name, a prompt, an eval.
- Whether you can add, remove or reorder the boxes.
- What an eval may check: length, wording, a fact, a rule.
- How many reworks a failing box gets, and what you see.
Three and four would change what an eval can hold, and neither is published. The three tests survive either answer, because each one asks whether a person can mark the box.
Write one eval per box before your first session. A standard you cannot fail is a preference.
Questions people ask next?
What are acceptance criteria?
The test a piece of work has to pass before it counts as finished, written before the work starts. Splunk’s September 2026 guide on evaluating AI systems begins there: define clear success criteria, then build the test sets.
How do you evaluate AI generated content?
Ask questions with a yes or no answer. Clarivate’s May 2025 guide uses two: does the response address the query, and does the source material support the answer. Someone who did not attend can answer both.
Can an eval check how a box sounds?
Not stated. Whether an eval can test length, wording, the presence of a fact, or a rule you typed is the third open question above. For now, keep the rule to something a reader can point at.
Ember Console is a session-report builder for coaches, in closed beta: https://ember.anyeads.com/. There is nothing to sign up for yet.
Ember