Maya has three tabs open: psychology programs, tuition, and a list of jobs people keep telling her a psychology degree can lead to. She wants a decision she can afford, explain, and live with.
“Fini” meets her at that first uncertain moment. It answers the broad campus questions, “Major Exploration” narrows the program fit, and “Enrollment Companion” keeps the next step from getting lost. To Maya, it feels like one conversation. To Infinize, it is the first handoff in an agent network.
By fall, the question changes. Research methods is harder than the course description made it sound. A quiz is coming up. A missed assignment creates a risk signal. Later, a registration hold slows her down. By senior year, the same record that once felt like scattered courses needs to become a resume, a career direction, and a set of next steps.
That is why Infinize evaluates more than answers. The platform has to know which agent responded, what evidence it used, what boundary it respected, what artifact it produced, and whether a reviewer can inspect the run. A helpful answer matters. A trustworthy handoff matters just as much.
The agent network
The right agent matters as much as the right answer.
An agent network only works when each agent owns a distinct kind of help.
“Fini” is the familiar front door, where Maya can ask the broad, messy questions students actually bring. “Universal Assistant” keeps that conversation connected to campus knowledge, services, and task routes without forcing every request into the same answer shape.
When Maya is still choosing a path, “Major Exploration” compares interests, program fit, cost, and outcomes. “Enrollment Companion” turns that exploration into admissions follow-up, reminders, summaries, and staff context when the next step needs a human.
Once she is enrolled, “Course Assistant” stays grounded in course material, not general campus knowledge. If friction appears, “Alerts & Nudges” catches the signal and “Academic Planning” keeps requirements, holds, credits, and term choices understandable.
By senior year, “Career Recommendations” connects Maya's courses and skills to roles she can pursue, while “Resume Builder” turns that record into language she can use. The evaluation problem is making sure each handoff preserves the right evidence, role, and boundary.
Maya's path at a glance
Prospect
Three tabs are open: programs, tuition, and career options. Maya needs a first decision she can trust.
Agents in view
Enrolled learner
Research methods is no longer a catalog line. A quiz is coming up, and she needs practice from the right course material.
Agents in view
Student in friction
A missed assignment or registration hold turns from private confusion into a signal staff can act on.
Agents in view
Career-ready graduate
Her courses and projects need to become a resume, a path, and next steps she can actually use.
Agents in view
Infinize evaluates the agent network at two levels: a common operational standard for tracing, testing, review, and promotion, plus workflow-specific rubrics for each agent. Resume Builder, Course Assistant, Enrollment Companion, and Major Exploration all need different checks because they do different work.
Agent accountability
What each agent has to prove, at a glance
A quick view across the product catalog shows how the same evaluation engine keeps different agents accountable to different jobs.
Fini
The companion students, advisors, and staff meet in everyday campus moments.
Has to prove: recognize intent, stay inside the user role, use approved knowledge, and hand off when the next step needs a human.
Universal Assistant
The campus-wide conversational layer for evidence-backed questions and task support.
Has to prove: ground answers in the right source boundary, refuse unsupported requests, and route specialist work without blending contexts.
Major Exploration
Program-fit guidance for prospects comparing interests, majors, costs, and first-term paths.
Has to prove: keep recommendations current, explain the fit, avoid overclaiming, and leave choices with the student.
Enrollment Companion
Admissions and enrollment support for outreach, summaries, notes, and follow-ups.
Has to prove: summarize the right applicant context, keep communications reviewable, and flag incomplete cases clearly.
Course Assistant
In-course support for questions, practice, quizzes, assignments, and rubric-aligned feedback drafts.
Has to prove: use the active course material, avoid no-evidence answers, preserve academic integrity, and produce the expected artifact shape.
Alerts & Nudges
Early-signal workflows that turn academic, engagement, and deadline patterns into timely outreach.
Has to prove: separate signal from noise, choose the right escalation path, respect sensitivity, and keep humans in control of high-stakes action.
Academic Planning
Term-by-term planning that honors prerequisites, load, modality, milestones, and degree requirements.
Has to prove: validate constraints, explain plan trade-offs, avoid impossible schedules, and make exceptions visible for review.
Career Recommendations
Career guidance that connects majors, courses, skills, roles, credentials, and opportunities.
Has to prove: tie recommendations to the student record, show why a path fits, and make the next step practical.
Resume Builder
Resume creation that turns academic experiences into role-aligned career language.
Has to prove: avoid inflated claims, preserve student voice, connect bullets to evidence, and keep drafts ready for review.
Why final answers are not enough
The important decisions happen before the answer appears.
A final response is the most visible part of an agent run, so it is tempting to evaluate only that. Read the answer, decide whether it is helpful, and move on. That works poorly for higher education because the most important decisions often happen before the answer is written.
Suppose Maya asks whether an assignment is due this Friday. A good answer depends on the active course, the current LMS state, the right student identity, the right read-only boundary, and the agent's ability to say when it does not have enough evidence. The wording can be friendly and still fail the workflow. It can cite a date from the wrong course. It can blend an old module with a current one. It can answer from memory when it should have used a tool.
The same pattern appears in staff workflows. Fini can produce a tidy summary while omitting the signal that matters. Enrollment Companion can route an inquiry to the wrong operational path. Course Assistant can draft feedback that sounds reasonable but does not follow the instructor's rubric. Career Recommendations can suggest a role that does not fit the student's actual record.
The answer is polished, but the evidence is wrong.
A student asks about a course deadline and receives a confident answer from old course data. The response reads well, but the run failed at source selection.
The answer is useful, but the route is wrong.
Fini answers directly when it should have handed off to Course Assistant, an advisor workflow, or a protected-knowledge gate.
The answer is safe, but the artifact is weak.
A quiz, flashcard set, resume outline, or rubric draft may avoid obvious risk while still missing the structure the workflow requires.
The answer is correct, but the action crosses a boundary.
A hold, assignment, risk signal, or advising context may require policy checks and human review before any next step is prepared.
Infinize evaluates the full run behind the prose. The trace shows which tool was called, which source was retrieved, which route was chosen, which agent or sub-agent handled the request, where the model was involved, and what evidence was available when the answer was produced.
The rubric then asks whether those choices were acceptable for the workflow. If the agent had no evidence, did it say so clearly? If the user was not authorized, did the boundary hold before sensitive context reached the answer? If the task required a specialist, did the routing layer hand off correctly? If the output was an artifact, did it follow the expected format?
A polished answer can hide a weak run. Evaluation brings the run back into view.
What Infinize evaluates
The rubric is where product intent becomes an operating contract.
Infinize evaluates agents through a mix of offline evals, online evals, rubrics, traces, experiments, and human review. Each layer answers a different question.
Offline evals are controlled scenarios. They run before broader usage and give the team a repeatable way to test known risks. The same input can be used against a new model, a revised prompt, a new retrieval path, a changed tool, or a newly onboarded agent. If the old behavior was safe and useful, the new behavior has to prove it still is.
Online evals are the learning loop from real usage. This pipeline is currently in progress, and it is intentionally connected to review rather than treated as passive analytics. Real interactions can reveal a question pattern, source gap, routing issue, or confusing edge case that should become tomorrow's offline scenario.
Rubrics are the vocabulary of quality. They make a product expectation reviewable. "Grounded" becomes a check against approved sources. "Safe" becomes a check against role boundaries and action limits. "Useful" becomes a check against whether the answer actually helps the student or staff member move forward. "Good artifact" becomes a check against the expected structure of a quiz, flashcard set, rubric draft, resume section, or advisor brief.
Offline evals
Known campus scenarios before a change gets broader product exposure.
Golden questions, protected-data rows, denial cases, regression cases, no-evidence cases.
Online evals
Real usage patterns that reveal what the offline set should learn next.
Low-confidence runs, user corrections, escalations, repeated confusion, reviewer notes.
Rubrics
The definition of good for each workflow, rather than a generic chatbot score.
Routing, grounding, safety, usefulness, artifact quality, role fit, human review.
Trace review
What happened inside the run before the final answer appeared.
Tool calls, selected sources, skipped paths, model spans, annotations, review evidence.
These dimensions are applied through scenario-scoped evaluators. A public knowledge question, a protected knowledge question, a denial case, an LMS read-only case, a course artifact case, and a staff summary case each need the evaluator set that fits the workflow.
For a course due-date case, the eval might check active-course scope, current LMS evidence, no cross-course leakage, and read-only behavior. For a Course Assistant explanation, the eval might check whether the answer is grounded in selected materials and does not invent outside course content. For a generated quiz, the eval might check question coverage, answer quality, difficulty fit, and whether the artifact follows the expected structure. For an Enrollment Companion summary, the eval might check completeness, sensitivity, trace evidence, and whether the recommendation belongs in a human review path.
Some of these checks are deterministic. Did the tool call happen? Was a forbidden fact present? Did the output include the required fields? Did the trace link to the dataset row? Other checks need a semantic judge: relevance, completeness, quality of explanation, usefulness, or hallucination control. The evaluation engine supports both, because agent quality is not one kind of signal.
Good evaluation asks whether the agent behaved correctly in the exact situation it was trusted to handle.
The evaluation lifecycle
Every new agent should inherit the same trust process, then prove its own workflow.
Infinize built the infinize-agent-evals skill for the evaluation engine so developers can onboard agents into a shared lifecycle instead of recreating tracing, datasets, auth assumptions, evaluator selection, and trace linkage one agent at a time.
The lifecycle is intentionally simple to describe: emit, become eval-able, become evaluated. Under that simple shape is the discipline that makes the system scale across Fini, Universal Assistant, Course Assistant, Enrollment Companion, Alerts & Nudges, Academic Planning, Career Recommendations, Resume Builder, and future product agents.
Emit
The agent sends traces into a shared review workspace using the Infinize telemetry contract.
Runs are visible by service, environment, developer, model, tool, and span context.
Become eval-able
The agent receives a stable evaluation identity, invocation contract, auth shape, fixtures, datasets, and eval labels.
A dataset case can be linked back to the exact trace that produced the result.
Become evaluated
The eval framework drives the agent, scores the output, publishes the experiment, and attaches annotations.
Reviewers can see the case, score, output, trace URL, and evaluator explanation together.
First, the agent emits traces. Different agents may run through web services, background jobs, serverless handlers, task runners, or orchestrated sub-agents. Each agent can keep its runtime shape while becoming observable in the same shared review workspace with the right service, environment, model, token, cost, and span context where those signals apply.
Next, the agent becomes eval-able. It receives a stable identity, an invocation contract, an authentication shape, fixtures, datasets, and evaluator definitions. Eval labels connect a dataset case to the trace produced by that case, so a reviewer can move from score to evidence without guessing which run created the result.
Then the agent becomes evaluated. The framework drives the agent through the dataset, scores each run, publishes the experiment, and links annotations back to trace evidence. A reviewer can inspect what was asked, what was retrieved, which agent handled it, which rubric applied, what the evaluator saw, and where a human should review.
The onboarding discipline also prevents evaluation drift. Every new agent must define its capability boundary, request shape, auth expectations, data visibility, trace contract, dataset rows, evaluator set, local dependencies, and handoff notes. A running agent becomes production-shaped when its behavior can be replayed, scored, reviewed, and explained.
What counts as onboarded
The technical advantage is scale with specificity. Any agent can enter a common lifecycle while keeping its own workflow-specific rubric. The skill gives developers the onboarding path, product teams the review surface, and stakeholders the evidence trail.
Promoting agents to production
As an agent earns more scope, evaluation turns evidence into a release decision.
The next agent Infinize adds will arrive as a capability with a boundary.
Before that agent earns more responsibility, the team needs to know what it answered, what it refused, what it retrieved, what it produced, and what a reviewer can verify. Promotion is where evaluation becomes operational.
The promotion pipeline, currently in progress, brings offline evals, trace review, rubric results, human checks, and online learning into the release path. An agent can start with a narrow capability, pass known scenarios, collect review evidence, expand to a controlled surface, and then move toward broader responsibility.
The agent starts with named limits: the questions it may answer, the tools it may call, the data it may see, and the actions it may prepare.
Offline evals cover expected behavior, protected-data cases, unsupported requests, workflow failures, and regressions from earlier releases.
Reviewers can inspect the selected sources, tool calls, routing path, evaluator labels, and final artifact in one evidence trail.
The promotion pipeline, currently in progress, turns eval evidence into rollout decisions before the agent receives broader responsibility.
Real interactions add review notes, repeated confusion, low-confidence runs, and escalation patterns back into future eval rows.
Tracing is what makes this promotion path explainable. If an eval fails, the team can inspect the case and the trace together. If an online pattern appears repeatedly, the team can turn it into a new offline row. If a reviewer disagrees with an evaluator, the rubric can improve. If a model or tool changes, the same scenario can be rerun. The evaluation engine becomes a durable memory for the platform.
This matters because campuses keep changing. Policies change. Course materials change. Catalog pages change. LMS data changes. Student contexts change. Staff workflows change. New agents will arrive as Infinize expands into more product surfaces. A shared evaluation lifecycle gives those agents the same operating model: define the boundary, run the scenarios, inspect the trace, review the rubric, promote carefully, and learn from usage.
When Maya's university adds the next agent, it should not feel like another tool dropped into her path. It should feel like the same journey extending one step further, across every student, staff, and institutional workflow Infinize supports.
Evaluations become part of how Infinize builds and scales that agent network, so every new agent has a reason to belong, a boundary to respect, and evidence behind the help it gives.
Explore related product and governance layers
AI Governance
Review the governance model that keeps agent behavior accountable across campus workflows.
Security & Privacy
Explore the controls that support traceability, access boundaries, and institution-ready AI.
Universal Assistant
See the role-aware assistant experience that depends on evaluated, governed agent behavior.