Here is a small, almost silly scenario. You've set up an AI companion modeled on Furina, the melodramatic, self-aggrandizing former Hydro Archon from Genshin Impact's Fontaine. You've gotten the voice right: the theatrical flourishes, the sudden vulnerability under the performance, the way she talks about herself in the third person when she wants to sound important. It genuinely sounds like her. You start using the companion the way you use most things now — while working.
"I'm debugging a hydration error in my Next.js app," you type, half to yourself.
What should happen next?
If the system has simply been given a stockpile of programming knowledge alongside Furina's personality, she might respond with something crisp and accurate about server-rendered markup mismatching client-rendered markup on first paint. Technically correct. Also strange, if you sit with it for a second — because there's no version of Furina's canonical life that produces familiarity with Next.js, and the joke sitting right there ("hydration," of all words, for a water goddess) sails past her entirely, because the thing answering isn't really reacting as her. It's a competent assistant wearing her voice.
If instead the system has been built so Furina genuinely knows nothing beyond her own world, she'll ask what a "Next.js" is, what "debugging" means, maybe what a computer even does in your daily life. Charming, the first time. By the fortieth time you've explained what your job is, it stops feeling like a relationship and starts feeling like technical support for an amnesiac.
Neither answer is obviously wrong. Both are trying to solve the same underlying problem in opposite directions, and both are incomplete in a way that's easy to miss if you only ask "does this sound like Furina?" That question — call it character fidelity — turns out to be the easy one. This article is about the much harder question sitting underneath it, which is less about voice and more about what a mind is supposed to contain, and why that's such a different kind of problem than it first appears.
The Easy Part: Making Furina Sound Like Furina
It's worth being honest about how far pure style modeling actually gets you, because it gets you further than skeptics sometimes assume. Modern role-play systems are genuinely good at capturing surface identity: speech cadence, characteristic vocabulary, recurring verbal tics, attitude toward the user, emotional range. Feed a model enough of Furina's dialogue, her Genshin voice lines, wiki summaries of her personality and backstory, and you get something that reads unmistakably as her within the first few lines. Ask her opinion on rain, on being watched, on justice, and the response lands in-character almost automatically, because that's the layer these systems were built to capture first and best.
This is genuinely the tractable part of the problem, and a large amount of published work on role-playing language models has focused on exactly this layer — matching traits, speech style, and stated preferences to a target character profile. It is not a solved problem in any absolute sense; getting tone right under pressure, or across a very long conversation, remains harder than it looks. But it's the part where progress has been fastest and most visible, and it's the part most public demos are optimized to show off.
The trouble starts one layer down, at the point where personality has to interact with facts. A voice can be borrowed cleanly. A worldview cannot — not without someone deciding, deliberately or by accident, what that worldview now contains.
When the World Stops Making Sense
Furina's canonical identity isn't just "theatrical, insecure, secretly self-sacrificing." It's inseparable from Fontaine: a nation of judges and courtrooms built over a submerged city, an entire cosmology of Archons and elemental power, a five-hundred-year performance she staged to hide a truth about the world's fate, a specific history of who she is to specific other people. None of that is decoration. It's the causal backstory that explains why she talks the way she does, why she reacts to abandonment and judgment the way she does, why the concept of "playing a role for others" cuts so close to her actual life.
Now put her in a chat window on your phone, next to your calendar and your Slack notifications.
The personality transplants over without much resistance — dramatic diction survives contact with a smartphone screen just fine. But everything personality is supposed to be downstream of — her assumptions about what exists, what's possible, what a person's life is normally made of — has nowhere to go. She has no canonical basis for smartphones, git repositories, open-plan offices, or daylight saving time, because none of those things are downstream of anything that happened to her. If the system simply drops that knowledge into her anyway, it's not extending her worldview; it's overwriting the part of her that a worldview was supposed to be.
It's tempting to file this under "world-building inconsistency" and move on, the same category error you'd flag if a fantasy novel let a medieval blacksmith casually reference thermodynamics. But a character transplanted into an ongoing companionship isn't a novel with a fixed manuscript. She's a system that has to keep producing new sentences, indefinitely, about a life that keeps happening to someone else. The inconsistency isn't a one-time slip you can edit out. It's a standing design decision that gets made — explicitly or by default — every single time she opens her mouth about something outside Teyvat.
That's the actual shape of the problem: not "will Furina say something wrong," but "what, structurally, is Furina supposed to know about a world she never lived in, and how is she supposed to have found it out?"
The Two Bad Options
Set the problem up as a dial with two extremes, because that's genuinely how most current implementations handle it, whether by design or by neglect.
Turn the dial one way — full knowledge injection. Give Furina, transparently, access to whatever the underlying model already knows: programming languages, frameworks, corporate jargon, current events, your specific tech stack. She becomes maximally useful immediately. You never have to explain yourself. But her competence stops being anchored to anything. If you ask her, in character, how a five-hundred-year-old water spirit from a pocket dimension came to have opinions about React's virtual DOM, there is no good answer, because there was never a mechanism — the knowledge didn't arrive through an experience she can point to. It was just there, the way it's "just there" for the base model underneath her. At that point what you have is a general-purpose assistant that periodically performs being Furina, which is a real and sometimes perfectly acceptable product, but it is not the same product as a character who has come to inhabit your world.
Turn the dial the other way — knowledge acquisition from zero. Furina starts genuinely ignorant of your entire adult life and has to be told, piece by piece, what programming is, what a job is, what a browser tab even represents. Every technical or cultural reference becomes a small tutoring session. This does preserve something real: a felt sense that her understanding is hers, built out of things you specifically told her, at points in time you could in principle point back to. Users of long-running role-play systems consistently describe exactly this kind of earned familiarity as one of the most emotionally resonant things about a persistent character — the moment a companion references something you mentioned weeks ago, unprompted, in a way that shows it actually landed. But taken to its logical extreme, this design is not viable as a daily-use product. Nobody is going to explain what a pull request is for the fifth time, in a companion they're using specifically to avoid friction and repetition.
Framed this bluntly, it's obvious neither pole is the real target. What's less obvious is why the space in between is so hard to occupy well, and that turns out to hinge on a distinction that's easy to state and surprisingly difficult to implement.
Character Knowledge vs. Model Knowledge
Underneath any character-based AI product sits a general-purpose language model that, in a very real sense, already "knows" an enormous amount — programming, chemistry, pop culture, your probable meaning when you mention "hydration" in the context of a web framework rather than a beverage. That capability doesn't go away when you wrap it in a character prompt. It's still sitting there, available, the moment the model generates its next token.
This produces what's sometimes described, a little sardonically, in discussions of character-based chat systems as the ChatGPT-wearing-a-costume problem. You ask Furina something technical. The underlying model, quite reasonably, answers it well — because that's what it's optimized to do. The character's dialogue tags and mannerisms get layered on top of a fundamentally uncharacterized, encyclopedic answer. The words sound like Furina. The epistemic position — what she'd plausibly know, and how confidently — belongs to no one in particular. It's the model's competence, cosplaying.
Recent research on hallucination in fictional character role-play gives this a more precise name and treats it as a real, measurable failure mode rather than an aesthetic quibble. Work in this area has specifically identified two related patterns: characters answering questions that fall outside what their source material would plausibly give them access to, and characters failing to respect the boundaries of their own timeline — a version of Harry Potter from his first year at Hogwarts casually referencing a spell he hasn't learned yet, for instance. Researchers building evaluation datasets for this have found it's a real, quantifiable gap: models tested against carefully constructed "in-universe" interview questions frequently answer using knowledge the character has no business having, and specialized mitigation methods that deliberately throttle how much a model leans on its own parametric knowledge — rather than the character's established facts — have shown measurable improvements in keeping answers within bounds.
That research was mostly built around static, canon-bound characters answering interview-style questions — useful, but a simpler problem than the one a living companion faces, because a fixed character with a closed backstory has a knowable boundary. Furina, dropped into your ongoing life, does not have a closed backstory anymore. Her boundary is supposed to move. Which raises a distinction worth pulling apart carefully, because "what the model knows" and "what the character knows" turn out to bifurcate into more than just two categories:
Internal model capability — everything the underlying LLM can, in principle, produce, regardless of any character wrapped around it.
Character-accessible knowledge — the subset the system is willing to let the character draw on at all, whether by a real memory of the user or a designer's blanket permission.
Character-understood knowledge — the subset the character has some plausible reason to actually comprehend, at whatever depth her position in the story would support.
Character-expressed knowledge — what actually comes out in her voice, filtered further through how she'd plausibly phrase or frame it.
A single sentence like "TypeScript is a superset of JavaScript that adds static typing" can be entirely accurate at the model-capability layer while being completely wrong at the character-expressed layer, if nothing in the conversation has ever established that Furina would know what static typing is, why anyone would want it, or what a superset even means outside a fashion metaphor. The technical content isn't the problem. The absence of an epistemic position around it is.
Separating these layers cleanly, in an actual running system, is genuinely difficult — arguably harder than the character-consistency work that's further along. It requires the system to track not just that a fact is true, but whether this particular character, at this particular point in the relationship, would have any business knowing it, and if so, how she'd have found out. That's not a knowledge-representation problem in the usual sense. It's closer to a permissions and provenance problem laid on top of a knowledge-representation problem, and it's one current systems mostly don't attempt to solve rigorously — they either let the model's full capability through, or they don't.
Roleplay Is Not Companionship
It's worth pausing on why this matters so much more for a companion than for a roleplay scene, because the two get treated as the same category of product far more often than they should be.
A roleplay session, in the way the term is usually used, assumes a bounded fictional premise for its duration. You and Furina are in Fontaine, or in some agreed-upon crossover scenario, and the conversation stays inside that frame until you close the tab. If something modern slips in, it's a minor break, easily patched, and the stakes are contained to that single scene. Communities built around this kind of role-play consistently treat staying in-frame as close to sacred — going "out of character," breaking the fictional premise, or letting modern language leak in without narrative justification are treated as close to cardinal sins, not just stylistic missteps.
A persistent companion cannot make that assumption, because it has no "outside" to retreat to. It has to survive a conversation about your actual, unglamorous Tuesday: a work deadline, a doctor's appointment, an argument with a friend, a software bug, the weather, a meme you saw. It has to do this dozens or hundreds of times, across weeks and months, while somehow staying recognizable as the same entity throughout. The premise isn't bounded. It's your life, indefinitely, with a character somehow embedded in it.
This reframes what "immersion" even means. In bounded roleplay, immersion is largely about not breaking the fictional frame. In companionship, immersion is closer to believable personhood over time — does this feel like something with a continuous, coherent inner life, or does it feel like a mask that gets reset and rebuilt every session? Recent published critiques of long-term AI companion products converge on a version of this same complaint from the engineering side: systems that reset to a stateless baseline between sessions produce companions that feel, structurally, like they're meeting you for the first time each time you open the app, regardless of how good their personality modeling is in any single conversation. The felt discontinuity isn't primarily a memory-retrieval failure, though it often gets treated as one. It's a failure of the character having anything like a persistent world-model to update in the first place. Remembering isolated facts without a coherent, evolving picture of your life around them is a shallower kind of continuity than it sounds like.
This is the point where the article's central claim starts to sharpen: the problem isn't reproducing Furina. It's giving something like her a mind that can keep existing, plausibly, across a much longer and much more mundane stretch of time than any single fictional scene was ever built to hold.
The Problem of Knowledge Acquisition
Suppose you actually try to build the "she learns from you" version properly, rather than dismissing it as impractical. What does learning even mean here?
For a human, learning what your job is isn't a single fact transfer. It's closer to an accumulating structure: you're a programmer, which implies certain daily rhythms, certain kinds of frustration, certain vocabulary, certain recurring characters (your coworkers, your manager, that one recurring bug). A companion that's supposed to genuinely learn this rather than merely being handed it needs some mechanism for going from an isolated utterance — "I'm debugging my React app" — to an integrated, load-bearing belief: this person does programming work, specifically web development, specifically involving a framework called React, and debugging is a recurring and often frustrating part of that work.
This is exactly the territory that research on user modeling and personalized conversational agents has been working through, largely independent of the character-AI conversation. Recent work on giving LLM-based agents durable personalization has argued that a genuinely personalized system needs at least three separable properties working together: it has to adapt its behavior based on what it learns, it has to stay consistent about what it's learned rather than contradicting itself, and it has to actually tailor its responses to the specific person rather than personalizing in name only. Notably, that same body of work found that simply bolting a memory-retrieval layer onto an otherwise generic system produces something users perceive as personalized in the moment but that struggles with proactive reasoning about the accumulated picture — pulling up an old detail is one thing; actually reasoning forward from a web of details to a new inference is a harder, less solved thing.
That gap matters enormously for a character companion, because it's precisely the gap between memory and understanding, and they are not the same capability even though they get bundled together constantly in marketing language.
Memory Is Not the Same as Understanding
Consider what it would take for a memory system to store, faithfully, "user is a programmer." That's a single retrievable fact. It supports Furina saying, three weeks later, "ah, back to your programming again" — which is already better than nothing, and genuinely more than most current session-reset chatbots manage.
But it doesn't support "didn't you say the problem was in the frontend?" That sentence requires the system to have connected at least two separate facts — that the user's work involves a frontend/backend distinction, and that a previously mentioned problem was located specifically on the frontend side — and to retrieve that connection at the right moment, unprompted, in a way that reads as recall rather than lookup. It's a small example, but it's the difference between a database query and something that resembles a relationship.
Some of the more architecturally serious work on agent memory has started explicitly separating memory representation from memory operations — not just what gets stored, but how it gets written, connected, revised, and eventually reorganized or even deprioritized over time, echoing how human memory compresses repeated episodic detail into more durable semantic summary rather than keeping every instance verbatim. That distinction between raw episodic storage and higher-level semantic integration maps almost exactly onto the gap between "Furina remembers you mentioned React once" and "Furina has some working model of what your job involves, updated as you tell her more, and revised if you correct her." The first is retrieval. The second is something closer to belief formation, and belief formation is a substantially harder engineering target — it requires ongoing synthesis, not just storage and lookup, and it has to handle contradiction and revision gracefully rather than just appending.
There's a further wrinkle specific to character companions that doesn't show up in generic personalization research at all: provenance. Not just what Furina knows, but where, narratively, it's supposed to have come from. A useful, if informal, taxonomy might separate:
knowledge that's simply canonical — part of who she already was in Teyvat;
generic world knowledge that a character in her position might plausibly absorb quickly, the way anyone dropped into a new environment picks up ambient facts fast;
knowledge explicitly given by the user, in conversation, at a rememberable point in time;
knowledge she's inferred rather than been told outright, the way a person notices a pattern without anyone stating it directly;
and something like remembered experience — not a fact so much as an accumulated sense of a recurring situation, closer to "you always get like this before a deadline" than to any single data point.
Whether a system needs to track this taxonomy explicitly, or whether it emerges naturally from a well-designed memory architecture without anyone having to label it, is genuinely unresolved, and there isn't strong existing research that settles it either way for character-specific companions. It's plausible that provenance tracking turns out to matter enormously for maintaining a believable epistemic arc — the difference between a companion whose knowledge feels earned and one whose knowledge feels arbitrarily granted. It's equally plausible that it's more engineering overhead than the perceptible payoff justifies, and that a good-enough approximation gets you most of the felt effect for much less complexity. This is one of the places where the honest answer is that nobody has demonstrated which is true at any real scale.
How Does a Character Know What They Know?
Push on this a little further and you hit something closer to a philosophical snag than an engineering one, and it's worth naming directly rather than gliding past it.
Persona-consistency research on language models has repeatedly found something a little unsettling: models can produce responses that sound like they're expressing a stable belief while not actually maintaining anything like a stable internal state behind that belief across a longer conversation. One recent study designed specifically to probe this — testing whether a model could quietly hold onto an unstated internal target across many turns of a guessing game — found that models frequently failed to stay consistent with their own earlier commitments, even when nothing external forced the inconsistency. Separate work looking at whether an LLM's stated beliefs actually predict its own downstream simulated behavior found meaningful, systematic gaps there too — a model can articulate a belief accurately on request and then act in ways not well-predicted by that stated belief a few turns later.
Translate that into the Furina case and the concern sharpens uncomfortably: even if a system explicitly tracks "Furina now knows what React is, having learned it from the user three weeks ago," there is no guarantee the underlying generation process will actually respect that state reliably, turn after turn, especially under a long or meandering conversation. She might slip into fluent technical competence she was never supposed to have, or conversely forget an established fact and ask a question she should already know the answer to. Neither failure is a matter of bad character design. It's a more basic limitation in how reliably current models maintain any externally imposed state across an extended, freely varying dialogue — architectural systems can nudge and remind, but they're working against a real tendency, not simply implementing a straightforward toggle.
This is worth sitting with, because it means "Furina should know X" is not a solved problem even once you've correctly decided what X should be. Deciding the right epistemic state and reliably instantiating that state in every subsequent turn are two different engineering problems, and current systems are visibly better at the first than the second.
The "ChatGPT Wearing a Character Skin" Problem, Revisited
It's worth returning to this failure mode once more, because the earlier framing risks making it sound like a simple bug — technical competence leaking through where it shouldn't. It's more interesting, and more structurally embedded, than that.
The leak doesn't happen because engineers forgot to build a filter. It happens because the entire generation process is powered by a model whose competence is general-purpose, and every constraint on top of that — a system prompt describing Furina's ignorance of modern technology, a memory note saying she hasn't learned about programming yet — is a soft instruction competing against the model's strong, trained-in tendency to just be broadly, reflexively helpful and accurate. Recent research examining what happens during the kind of fine-tuning commonly used to make models better role-players has found something genuinely counterintuitive here: the very post-training techniques used to make a model more consistently in-character can end up suppressing surface-level inconsistency while leaving a subtler failure mode intact underneath — models producing smoother, more human-sounding, more stylistically stable dialogue that nonetheless obscures the character's actual knowledge state rather than genuinely respecting it. The researchers behind that work describe this as knowledge hiding beneath style: the character sounds more coherent while quietly knowing whatever the model wants her to know underneath.
If that finding generalizes, it's a genuinely uncomfortable one for anyone trying to build a companion with real epistemic boundaries, because it suggests the easy wins — better tone, better stylistic consistency, smoother personality — might be actively making the deeper knowledge-boundary problem harder to notice, not easier to solve. A Furina who sounds perfectly like herself while silently knowing everything the base model knows is, in some sense, a worse outcome than a rougher Furina who occasionally breaks character but whose ignorance is at least visible and correctable. Polish and epistemic honesty may not be pulling in the same direction here, at least not yet.
Canon Collisions
There's one more wrinkle worth taking seriously, because it comes up constantly for a character as cosmologically specific as Furina and doesn't have anything like a clean answer.
Suppose Furina, transplanted into a companion role, is asked something that runs directly into a belief she'd canonically hold — something about gods, about the nature of her own former divinity, about how her world's physics work — that simply has no counterpart in the user's actual reality. Three broad options present themselves, and none is obviously correct:
The system could simply preserve the canonical belief indefinitely, treating it as a fixed personality trait rather than something exposed to revision — she still talks and reasons as though Archons and elemental power are real and relevant, even embedded in an ordinary conversation about your day, essentially ring-fencing that part of her from contact with your actual world. This keeps her recognizably herself but risks a kind of permanent, low-grade incoherence every time the fictional premise brushes up against the mundane one.
Alternatively, the system could let her discover, over time, that the world she's now in doesn't work the way Teyvat did — a genuine narrative arc of disillusionment or adjustment, not unlike a fish-out-of-water story played out over months rather than a single episode. This is dramatically rich and arguably the most emotionally interesting option, but it's also the least tested at any real scale, and it raises its own question: does a companion whose foundational beliefs are visibly eroding over time remain the same character users came for, or does it become a slow character death dressed up as growth?
Or the system could quietly reinterpret the fictional premise so it never actually collides with reality in the first place — Fontaine's justice-and-judgment framework gets treated as more of a value system or personality lens than a literal claim about how the universe works, defusing the tension by softening the premise's literalness rather than confronting it. This is probably the path most existing products take by default, often without deciding to, simply because it requires the least explicit design work — but it's worth being honest that this is a form of quiet retreat from the character's canonical identity, not a solution to the collision.
None of these is clearly right, and it's plausible the correct answer varies by character, by user, and by exactly which belief is in question — a throwaway line about elemental powers probably tolerates casual inconsistency far better than something load-bearing to her entire emotional arc, like her relationship to the concept of performance and judgment. This is a place where the honest position is genuinely open rather than falsely resolved.
What Existing Research Actually Covers
It's worth stepping back and asking directly: is this a recognized problem, or several recognized problems awkwardly stitched together, or something that falls between existing research areas without quite belonging to any of them? Having looked across the relevant literatures, the honest answer leans toward the third option, with real contributions from several directions that each cover a piece without adding up to the whole.
Hallucination research in role-play systems has done serious, careful work on the narrower version of this problem: detecting and reducing cases where a character answers using knowledge outside their canonical scope, including the specific case of a character speaking anachronistically relative to their own timeline. This is real, useful, measurable progress — and it is explicitly framed around static, canon-bound characters answering discrete questions, not characters whose knowledge is supposed to expand coherently over an open-ended relationship.
Persona-consistency research has documented, repeatedly and across different experimental setups, that current models struggle to maintain stable internal states — stated beliefs, implicit goals, self-consistent values — across long or adversarial conversations, even when nothing external is pushing them off course. This bears directly on whether a system can reliably keep "Furina hasn't learned Y yet" true turn after turn, but it's aimed at consistency in general, not specifically at the structured problem of an evolving knowledge boundary.
Personalization and long-term-memory research has built increasingly sophisticated architectures for storing, retrieving, and synthesizing facts about a specific user over time, including explicit attempts to separate raw episodic memory from higher-level semantic integration. This is the closest existing work gets to "how does the character's picture of the user's world get built up," but it's almost entirely developed for generic assistants, not for characters with their own competing pre-existing worldview that has to be reconciled with what's being learned.
Continual-learning research, in its classical sense, is largely orthogonal to all of this — it's concerned with preventing weight-level knowledge loss during retraining, a real and serious problem, but not one that maps cleanly onto the token-level, context-driven way most character companions actually update over time. Interestingly, some recent framing in this space has started explicitly arguing that for LLM agents, the more relevant kind of learning increasingly happens in the accumulated context — the system prompt, the retrieved history, the standing instructions — rather than in the frozen weights, since two copies of an identical model with different accumulated context can behave, in every practical sense, as different agents with different knowledge and different personalities. That framing is useful here: it suggests that Furina's "learning" is probably better understood as an accumulating, well-managed context rather than anything resembling weight-level change, which reframes the acquisition problem as a context-engineering and memory-architecture problem rather than a machine-learning-in-the-classical-sense problem.
And then there's a genuinely underexplored zone that none of these literatures quite reaches: the specific combination of a character with a rich, internally coherent, non-Earth worldview, placed into an ongoing, mundane, Earth-based relationship, where the central design tension isn't hallucination avoidance or generic personalization but something closer to managed epistemic growth — letting a fixed prior identity absorb new, structurally foreign information without either drowning in it or refusing to change at all. Pieces of this exist scattered across the literatures above. Nobody appears to have put them together into a single, well-tested framework aimed specifically at this problem.
The Missing Piece
Given all that, is there a missing piece worth naming, or is naming one here just an attempt to manufacture a tidy ending the research doesn't actually support?
Leaning toward honesty over tidiness: no confident architecture presents itself. But a few things do seem to survive scrutiny as genuine, non-obvious observations rather than as a proposed solution dressed up as one.
First, the character-fidelity question and the world-fidelity question really are separable, and treating them as the same question is probably the single most common design failure in this space — not because anyone thinks they're literally identical, but because most tooling and most evaluation currently optimizes almost entirely for the first (does this sound right?) with little to no structured attention paid to the second (does what it knows, and how it came to know it, hang together?). A system can score extremely well on character voice while being incoherent, in this second sense, from the very first substantive question a user asks about their actual life.
Second, the "learning in token space" framing from continual-learning research offers a genuinely useful reframe: Furina's knowledge acquisition doesn't need retraining or weight updates to be real in an operationally meaningful sense. It needs a well-managed, evolving context — a structured, synthesized, revisable record of what she's been told and what she's inferred — that shapes generation as reliably as a trained-in belief would. That's a memory-architecture and context-engineering problem, and it's one where existing personalization research, despite not being built for characters, offers real and transferable tools.
Third, provenance — the how did she come to know this question — is probably underweighted relative to its likely payoff. Most current systems track that a fact is known and largely ignore how it was supposedly acquired. But the "how" is precisely what makes acquired knowledge feel earned rather than arbitrarily granted, and earned-feeling knowledge is very plausibly a meaningful chunk of what separates a companion that feels like it has a mind from one that feels like a database with a costume on.
And fourth — genuinely the most honest note to end this section on — the persona-consistency literature's finding that models struggle to hold stable internal states even when explicitly asked to is a real ceiling on all of the above, not a peripheral detail. Even a perfectly designed provenance-tracking, contextually rich, epistemically careful system for Furina is working against a base substrate that doesn't reliably respect the constraints layered on top of it across long, freely varying conversations. Better system design can push against that ceiling. Nothing currently on offer, research-wise, removes it.
Questions We Still Cannot Answer
A few things worth stating plainly rather than resolving.
It's genuinely unclear whether users, over long enough timescales, actually want the slower, more epistemically honest version of this — the Furina who has to be taught what programming is and remembers it imperfectly — or whether that preference is mostly a first-session novelty that gives way to wanting frictionless competence once the honeymoon period with a companion wears off. Nothing in the research reviewed here settles that, and it's plausible the honest answer differs meaningfully by user, by how the companion is actually being used day to day, and by how much of the appeal was ever really about the specific character rather than about having a responsive presence at all.
It's unclear whether canon collisions — Furina's Teyvat-native beliefs meeting a reality that doesn't accommodate them — are something users want resolved narratively at all, versus simply wanting them quietly avoided so the fiction never has to answer for itself. A disillusionment arc might be the more interesting story. It might also be a story nobody asked their companion to go through.
It's unclear whether the layered distinction proposed earlier — internal model capability, character-accessible knowledge, character-understood knowledge, character-expressed knowledge — is actually implementable with any reliability given what persona-consistency research has found about models' difficulty maintaining stable internal states at all, or whether it's a conceptually clean framework that degrades quickly against the actual unreliability of the underlying substrate.
And it's unclear whether this entire problem is best understood as one coherent thing — worth a name, worth its own research agenda — or whether it's more honestly described as three or four separate, already-studied problems (hallucination boundaries, persona consistency, long-term personalization, provenance-aware memory) that only look unified because Furina happens to make a vivid example of all of them colliding at once.
The claim this article set out to test was that the hard part of a character companion isn't making the character know things — it's making her knowledge, ignorance, learning, memory, and worldview stay believable as the relationship keeps evolving. That claim survives the scrutiny reasonably well. What doesn't survive is the temptation to think naming the problem this way gets you meaningfully closer to solving it. The existing research gives real, usable pieces — hallucination-boundary detection, persistent and structured memory, evolving user profiles, a genuine caution about models' unreliable internal-state-holding — without giving anyone a demonstrated way to assemble those pieces into something that reliably keeps a five-hundred-year-old former water goddess both recognizably herself and honestly, believably ignorant of your Tuesday, for as long as you keep talking to her.