Designing a usable system is not enough if one cannot measure, with rigour and systematicity, the extent to which those objectives have actually been achieved. Evaluation is not the final phase of a digital project: it is the continuous process that transforms design hypotheses into verified knowledge.

The pages dedicated to Esario's individual services — Usability, User Experience and Accessibility — describe the what and the why of each discipline. This page addresses the how: the tools, methods and metrics with which Esario measures the quality of interaction between people and digital systems, integrating objective and subjective approaches into a coherent and multidimensional evaluation process.

An integrated approach to evaluation

The evaluation of interaction quality between user and system is a field where different disciplines converge: cognitive ergonomics, experimental psychology, behavioural sciences, software engineering. This convergence produces extraordinary methodological richness and, at the same time, a complexity that requires orientation and expertise to be governed.

International research has made clear that no single tool is capable of capturing the quality of an experience in its entirety. Usability — in its dimension of effectiveness, efficiency and satisfaction — can be measured with objective performance metrics. But User Experience, with its emotional, aesthetic and contextual components, requires different, more nuanced instruments. And accessibility requires yet another level of analysis, one that integrates technical conformance with standards and the verification of the real experience of users who rely on assistive technologies.

Objective and subjective metrics are not alternatives — they are complementary. Their integration is the necessary condition for a reliable evaluation that measures experience in its entirety.

In practice, this means selecting for each project the combination of tools best suited to the context, the users and the evaluation objectives — avoiding both the superficiality of a single questionnaire and the dispersion of an unfocused approach.

Categories of metrics and tools

Standardised subjective scales

Standardised subjective scales are questionnaires validated by scientific research, designed to measure the user's perception across specific dimensions of usability and User Experience. They are the most widely used tool in UX evaluation thanks to their reliability, comparability across studies and ease of administration.

The System Usability Scale (SUS) provides a synthetic and robust measure of overall perceived usability. The User Experience Questionnaire (UEQ) assesses perceived quality across six dimensions covering both pragmatic and hedonic aspects. AttrakDiff separately measures pragmatic and hedonic qualities, making it possible to distinguish between perceived ease of use and the pleasure, stimulation and sense of identity the product conveys. The choice of scale is never arbitrary: it depends on the construct to be measured, the user population and the phase of the product lifecycle.

Post-task and post-session questionnaires

Administered immediately after the completion of a specific task or at the end of a test session, these tools capture the user's perception while it is fresh, reducing the memory biases that accumulate over time. The Single Ease Question (SEQ) measures perceived difficulty of the task just completed. The After-Scenario Questionnaire (ASQ) assesses satisfaction after a structured use scenario. The Computer System Usability Questionnaire (CSUQ) measures overall satisfaction at the end of a session. Their brevity makes them particularly suited to intensive test sessions, where participant fatigue is a critical variable to manage.

Cognitive and performance metrics

These are objective measurements of user behaviour during interaction: task completion time, success rate, number of errors, efficiency understood as the ratio between correct output and resources employed. They are the metrics most directly aligned with the definition of usability in ISO 9241-11 and constitute the quantitative foundation of any empirical evaluation.

To these is added the NASA Task Load Index (NASA-TLX), a tool originally developed for high-complexity operational contexts and today widely used in the evaluation of complex digital interfaces — particularly in the Finance and Bio-Medical sectors where the informational density of systems is often high and the management of cognitive load is a critical dimension of design.

Behavioural metrics

Collected passively through the analysis of system logs, behavioural data reflect the real behaviour of users in natural or semi-natural contexts, without requiring any interruption of the activity or any explicit declaration from the user. Clickstream analysis reveals actual navigation paths, sequences of actions, hesitation points and areas of confusion that no questionnaire would be able to detect with the same precision. Abandonment rates signal where the experience breaks down before the completion of the goal. Retention and frequency-of-use data measure engagement over time — one of the most reliable indicators of overall experience quality, since it reflects not the declared judgement but the real and repeated behaviour of the user.

Emotional metrics

To measure the user's emotional response beyond the simple judgement of usability, tools specifically designed to capture the affective dimension of experience offer a complementary and indispensable perspective. The Self-Assessment Manikin (SAM) evaluates valence, arousal and dominance through non-verbal visual representations, accessible to users of any cultural and linguistic background. The PANAS measures the prevalence of positive and negative affective states during and after interaction. PrEmo makes it possible to identify the specific emotions elicited by a product through a visual vocabulary of animated characters — particularly useful in contexts where verbalisation of emotions is difficult or culturally conditioned.

These metrics are especially relevant in the Hotels and E-Commerce sectors, where the user's emotional response is a critical determinant of conversion, loyalty and the perceived value of the service.

Longitudinal and retrospective methods

The quality of experience is not static: it transforms over time, with learning, with repeated use, with the progressive integration of the system into the user's daily routine. Longitudinal and retrospective methods are designed to capture this temporal dimension of experience, which point-in-time evaluations are unable to grasp.

The Experience Sampling Method (ESM) collects real-time assessments through periodic notifications, offering maximum contextual authenticity. The Day Reconstruction Method (DRM) asks users to reconstruct their day episode by episode, balancing accuracy and intrusiveness with a significantly lower burden. The UX Curve and iScale allow the evolution of experience to be visualised graphically over time, identifying the turning points that have determined shifts in the perception of the system.

Research has demonstrated that retrospective evaluations predict future user behaviour better than daily assessments, making them tools of great strategic value for organisations that want to understand not only initial adoption but loyalty over time.

The metrics table: an operational guide

The table below presents a selection of the most significant metrics and tools for the sectors in which Esario operates, showing category, type of data produced and primary domain of application.

Metric / Tool Category Domain Esario Sector
SUS – System Usability ScaleStandardised scalesOverall perceived usabilityAll sectors
UEQ – User Experience QuestionnaireStandardised scalesPerceived UX quality (pragmatic/hedonic)E-Commerce, Hotels
AttrakDiffStandardised scalesPragmatic and hedonic qualityE-Commerce, Hotels
SEQ – Single Ease QuestionPost-task questionnairesPerceived task difficultyAll sectors
CSUQPost-session questionnairesOverall satisfactionFinance, Bio-Medical
Task completion timeCognitive metricsEfficiencyAll sectors
Success rateCognitive metricsEffectivenessAll sectors
NASA-TLXCognitive metricsCognitive loadBio-Medical, Finance
Clickstream / Log analysisBehavioural metricsNavigation, efficiencyE-Commerce, Finance
Abandonment rateBehavioural metricsEngagement, usabilityE-Commerce, Hotels
SAM – Self-Assessment ManikinEmotional metricsValence, arousal, dominanceHotels, E-Commerce
PANASEmotional metricsPositive/negative emotionsBio-Medical, Hotels
PrEmoEmotional metricsProduct-elicited emotionsHotels, E-Commerce
ESM – Experience Sampling MethodLongitudinal methodsExperience in real contextHotels, Finance
UX CurveRetrospective methodsEvolution of UX over timeAll sectors
WCAG 2.1 / EN 301 549 AuditAccessibility evaluationStandards conformanceAll sectors

Evaluation and Artificial Intelligence

The integration of Artificial Intelligence into evaluation processes is significantly expanding Esario's capacity to produce deeper, faster and more comprehensive analyses.

In the automated analysis of behavioural data, AI models make it possible to process volumes of logs and interaction data that would be impossible to analyse manually — identifying recurring patterns of difficulty, systematic error sequences and correlations between behaviours and system characteristics that would escape traditional observation.

In accessibility evaluation, AI-assisted automated scanning tools make it possible to conduct conformance audits on complex digital systems with a coverage and speed impossible through manual methods alone, complementing the testing conducted with real users who employ assistive technologies.

In the analysis of sentiment and emotional responses, language models and multimodal analysis make it possible to extract experience quality signals from heterogeneous sources — reviews, support session transcripts, free-text feedback — building an emotional map of the experience that completes the data collected through standardised scales.

In the simulation of user journeys, Large Language Models can generate test scenarios, simulate the behaviour of users with different characteristics and competencies, and produce structured heuristic evaluations that accelerate the expert analysis phase without replacing it.

At Esario, AI tools do not replace methodological expertise: they amplify it, making it possible to bring greater analytical depth to every evaluation project, regardless of its scale and complexity.