Evaluation guide

Do not judge an AI front desk by how human it sounds.

Judge whether it preserves context, retrieves instead of guessing, completes an action, handles correction and knows when a person should take over.

Reproducible testPush the edge caseInspect operator result
Evaluation checklistOPERATING VIEW
Context

Does it remember what you already told it?

Test
Grounding & action

Does it retrieve facts and complete work?

Test
Human boundary

Does it know when to stop and hand over?

Test

Use the same six tests on every AI front desk.

These tests deliberately move beyond a clean greeting and a scripted FAQ.

1. Memory

Tell it a location or constraint, then continue without repeating it.

2. Retrieval

Ask for availability or a live fact that should come from a system rather than model memory.

3. Action

Move from enquiry into a booking, change, viewing or another observable outcome.

4. Correction

Change your mind and see whether the workflow updates instead of restarting.

5. Handoff

Ask for something that should require a person and inspect whether the context travels with it.

6. Operator visibility

Check whether the team can see completed work and exceptions without reading every conversation.

Use a scenario that has a clear end state.

A completed booking is easier to evaluate than a friendly but open-ended conversation.

REPRODUCIBLE TESTLive Trelinx demo · illustrative business data
Try this exact request

“I want a free trial this Saturday at Bishan.”

Checks the scheduleBooks the trialKeeps the record
Then change something:“Can I move it to next Thursday?”
Call TrelinxNo signup. Test a working demo agent.

The strongest result is boringly operational.

The AI does not need to impress the customer with intelligence. It needs to leave the workflow in the right state.

Evaluation resultILLUSTRATIVE WORKFLOW STATES

Outcome exists

There is a booking, update, link or another concrete next state.

Completed

Change is preserved

The update applies to the same object rather than starting over.

Completed

Exception is prepared

The human receives the context and only does the irreducible judgement.

Human needed

Run the same test against Trelinx and any alternative.

The category should compete on operational completion, correction and control.

Try Trelinx