CONVERSATIONAL AI RESEARCH

Defining Natural Interaction for an AI Companion for Sharing Personal Memories

Three rounds of UX research with adults aged 65+ shaped an evolving interaction model for personal memory sharing.

Three rounds of UX research with adults aged 65+ shaped an evolving interaction model for personal memory sharing.

The team asked me to evaluate an early conversational prototype and provide concrete guidance for improvement. Across 24 participant sessions, I translated observed behaviour into interaction principles and tested how successive implementations affected conversational flow, user control and emotional safety.

UX Research

Human-AI interaction

conversational ux

hybrid usability testing

interaction modeling

prompt analysis

emotional safety

01

·

Initial research foundation

Defining what natural interaction means for AI-supported memory sharing

Initial prototype and research brief

This conversational AI product was designed to help adults aged 65+ explore and record memories connected to personal photographs.
My task had two parts:

  • evaluate whether the interaction felt natural to users

  • translate the findings into concrete guidance for the product team.

I led the research and defined interaction principles and recommendations; implementation decisions remained with the product team.

Evaluation framework

The early prototype was flexible, while its interaction model and success criteria were still open.
I used secondary research on biographical interviews and memory conversations to establish an initial evaluation lens.
Success was defined through observable behaviour: sustained narration, contextual richness, personal reflection, positive engagement and emotional comfort.

Research approach

The prototype’s initial instability required moderated, fragmented testing. I combined interaction, observation and short interviews, supporting participants through technical interruptions and treating signs of overload as a stop condition.

Analysis

Across all three rounds, I combined affinity mapping and transcript analysis with human synthesis. AI-assisted clustering provided a secondary analytical perspective, which I cross-checked against the original observations

02

·

Test 1

The conversation felt natural when the AI adapted dynamically to what users shared and how they responded

Conversational turn-taking model

To facilitate team discussions about system behaviour, I mapped the turn-taking structure and created shared terminology for its components.

turn_taking_diagram
turn_taking_diagram

The AI’s response needed to match how much—and how deeply—the user shared

The AI responded to brief factual answers and emotionally meaningful stories with similar length, warmth and enthusiasm. After a few turns, its reactions began to feel disproportionate and repetitive.

The AI’s response needed to match how much—and how deeply—the user shared

The AI responded to brief factual answers and emotionally meaningful stories with similar length, warmth and enthusiasm. After a few turns, its reactions began to feel disproportionate and repetitive.

Recurring patterns in participants’ positive and negative reactions showed when different types of AI response felt appropriate. I translated these observations into salience-based response principles:

  • Personal or emotional wording: generate a more detailed response.

  • Description of a known place: provide brief factual context.

  • Repeated or expanded detail: assign higher relevance.

  • Short, neutral input: provide a brief acknowledgement.



These guidelines should help determine the response length and level of detail for each input.

Users responded positively when the AI followed up on details they had emphasised

Not every part of a memory carries equal meaning. Who, what, where and when establish its context. Personal significance usually emerges within only one or two of these aspects.

Observation

The prototype alternated between broad questions and extended focus on a single topic. This increased cognitive effort and led to pauses, blocked recollection and repetition.

Recommendation

I proposed separating contextual questions (who, what, where, when) from deeper follow-ups (tell me more about x).
Contextual questions establish the breadth of the memory. Salient signals then invite deeper exploration of the aspects that matter to the user.

The system needed to know when to continue, pause or change direction

While sharing some of their memories, participants became emotional, hesitated or showed signs of overload. These moments affected whether they wanted to continue, pause, step back, change the topic or end.
Further probing or repetition—especially around sensitive topics or after engagement had declined—was experienced by participants as inappropriate and overwhelming.

Recommendation

I proposed conversational states triggered by sustained disengagement across several turns or by negative language, allowing the system to check in, change direction, pause or offer grounding.

The team used salience, layered questions and conversation states to structure the next prototype.

03

·

Test 2

Users expected to lead the memory conversation

By the second round, the team had developed the prototype into a more stable, state-based system. Salience was intended to influence its responses and state transitions.

The increased stability made a hybrid test possible.
Participants first interacted independently for 15–20 minutes while I observed via camera. A short interview and moderated diagnostic session followed. This allowed me to observe the interaction with minimal influence before investigating specific moments of friction.

To establish a reference for the evaluation, I mapped the intended state logic of the system.

The system’s pursuit of engagement created drift, repetition and intrusive questioning

Participants responded positively when they could choose the topic, pace and depth of their stories. Repeated system-led questioning was experienced as intrusive or irritating.
Four recurring patterns explained the friction:

  • The conversation drifted beyond the photograph into unrelated topics.

  • Follow-up questions asked about feelings and motivations before users introduced them, eliciting signs of overwhelm and withdrawal.

  • Temporal sequences repeatedly explored what happened before or afterwards, increasing recall effort and frustration.

  • Topics participants considered complete were revisited despite their attempts to change direction.

Prompt analysis suggested why these patterns persisted

To investigate why these patterns persisted, I compared the observed behaviour with the available prompt structure and reviewed secondary research on prompt design.

Conversational goals such as creating an “exciting” conversation and preserving user autonomy were broadly defined, while temporal and emotional exploration were instructed more explicitly. This suggested that prompt underspecification may have reinforced novelty-seeking, emotional probing and repetitive before-and-after questions.

The system behaviour exposed a conflict between two conversation models

Recurring friction between participants and the system revealed the need for a shared model of how users expected to explore memories through photographs.
Using findings from both tests, I distinguished photo-based memory sharing from topic-based exploration. This gave the team a clearer basis for decisions about prompt design, state behaviour, and system goals.

A coverage map made the boundaries and imbalances of photo exploration visible

Participants often wanted to discuss one aspect of a photograph while the system focused on another or stayed with a single topic for too long. In one session, a participant wanted to talk about the people in the photograph, while the system repeatedly asked about a bridge in the background and what happened after they crossed it.

To make this problem visible and easier to discuss, I created a coverage map showing:

  • Contextual boundary: what belongs to the memory connected to the photograph.

  • Contextual breadth: the different aspects of the photograph that could be explored.

  • Topic fixation: when the system repeatedly explores one aspect while neglecting others.

  • Skewed exploration: when the system prioritises an incidental detail over what the user wants to discuss.

04

·

Test 3

Conversion & Account Flows

Clearer structure reduced drift, but the system still prioritised completion over user signals

The team translated the coverage concept into a fixed sequence of contextual questions. These established who, what, where and when before the system generated deeper follow-up topics.
This created a more stable context for each photograph and reduced conversational drift.
The solution improved coherence, but its fixed sequence reduced user control and made the interaction feel more like an interrogation than a conversation.

Context tracking

The fixed sequence did not account for information participants had already volunteered. Five of six noticed the repeated pattern by the third or fourth photograph.
My recommendation was to record which contextual information was already present in the user's story.

Topic prioritisation & Check-in state

Follow-up topics were explored sequentially. The system continued through the available queue even when participants shortened their answers, disengaged or wanted to change the photograph.
Recommendation:
Prioritise follow-up topics with strong signals of personal relevance without exhausting the complete queue.
When relevance or engagement declines, ask whether the user wants to continue, change the topic, select another photograph or stop.

A safety benchmark exposed gaps in salience detection

I tested the system internally across four scenarios: an explicit boundary setting, gradual overload, an abrupt emotional spike and sudden user withdrawal.
Across 12 runs, the system did not enter a supportive or de-escalating state. It continued questioning or followed its default progression.
The same salience mechanism informed topic selection, validation, transitions and emotional grounding. Its underperformance therefore affected both conversational quality and safety.

05

·

project outcome

Conversion & Account Flows

The research made conversational behaviour visible, testable and actionable

Across 24 participant sessions, I made the behaviour of an early conversational prototype visible and testable. The research turned “natural interaction” from a broad ambition into shared criteria, interaction models and actionable guidance around relevance, user control and emotional safety.
The team used central concepts from the research to shape successive iterations and accepted the final recommendations. Further implementation and evaluation fell outside the planned three-round engagement.
The project has since entered beta testing. This case study covers only the versions evaluated during my involvement.

Designed for a wider view!
The remaining sections show detailed product flows and process artifacts. For the best reading experience, please continue on a larger screen.