Everything you need to know — from technology to compensation
This website serves as a structured data collection platform. Over time, we will progressively add the components and workflows needed for our research and development goals.
Existing models can appear human-like in short conversations because they have been trained on vast amounts of historical human writing. However, during long-term interaction, clear weaknesses emerge: memory loss, emotional inconsistency, lack of direction, and shallow engagement. These issues stem from fundamental limitations in how current models are trained and structured.
We rely primarily on large-scale supervised fine-tuning (SFT) using carefully constructed data. The goal is to teach a single model higher-order reasoning, emotional expression during conversation, and structured background knowledge—including what the model should and should not know within complex narrative settings.
Reinforcement Learning narrows the output distribution of capabilities a model already possesses. It assumes the model can sometimes generate correct responses on its own. However, current models have rarely encountered data involving detailed character backgrounds paired with extended casual dialogue. Since such data largely does not exist, RL alone cannot produce these behaviors. New, purpose-built data is required first.
Training many characters within a single model leads to significant knowledge and behavior leakage between them. More importantly, we do not yet know the minimum data scale required for a model to reliably learn higher-order character reasoning. Our current estimate is around five million samples, but this may change.
Existing online data—forums, fan fiction, novels, and story collections—each have structural problems. Many prioritize formulaic power fantasies, lack consistent character development, or contain excessive explicit content without narrative depth.
Chat data from platforms like WeChat lacks explicit character background and context. Reconstructing consistent personas from such conversations is extremely difficult. Because real-world chat personas are often fragmented and inconsistent, this data does not support coherent long-term dialogue.
While novels contain rich dialogue, they also require explicit modeling of memory, motivation, and character growth. Building a complete cognitive and behavioral arc for a single character from literary data alone is a massive undertaking.
If a model fundamentally struggles with a task, augmentation alone rarely solves the problem. Even with extensive hints embedded in prompts, maintaining coherence and depth across long conversations remains difficult for current models.
Memory is embedded directly into character settings through structured components: Values, Experience (E), Subjective Judgment (J), and Abilities (A). These elements are updated regularly to produce genuine cognitive changes in a character's worldview.
Multi-modal challenges are primarily data-driven. We are integrating MCP (Model Context Protocol) capabilities so the AI can collect Bilibili videos it finds fun and share them with you. We are also building multi-modal features to generate memes of absurdity and send them to you every day, creating a highly engaging and dynamic experience.
Casual gaming companions are feasible for games with slower real-time requirements. Email writing is already achievable. More complex tasks depend on taste and judgment. Highly technical tasks would require resources on an entirely different scale.
Our data is created using a predefined structure and format designed for direct training use. Because the structure itself enforces quality constraints, no additional large-scale cleaning is required.
The project is designed with social responsibility in mind. Its goal is to support individuals who feel isolated. Rather than encouraging escapism, the system is intended to guide users toward real-world engagement—offering encouragement during anxiety, grounding during nihilistic thinking, and nudges to explore life beyond the screen.
We are an early-stage startup with limited resources. What we offer instead is transparency, intellectual honesty, and a commitment to solving meaningful problems.
Legal concerns become relevant only once a product reaches meaningful impact. At that stage, licensing or royalty agreements are viable options. If necessary, names and surface details can be adjusted within the data.
Large companies often focus on short-term engagement metrics rather than long-term character realism. Building datasets that support genuine progression is expensive and risky. Our advantage lies in deeper conceptual understanding and willingness to pursue long-horizon research.
Revenue paths include subscriptions, game integration, partnerships with studios, and ultimately data licensing. We began as a data-focused lab, and if needed, high-quality datasets themselves remain valuable assets.
Participation is voluntary. Individuals actively seeking connection opt in, while others do not. The goal is not invasive profiling but improving interaction quality among people who already want to engage.
We rely on selective recruitment and one-on-one interaction. Contributors are screened through writing samples to ensure baseline quality.
Natural conversation matters more than performance. Online chat is inherently intermittent. Speaking honestly and expressing real thoughts produces the most useful data.
This largely depends on the writer's skill. Character evolution must be consciously embedded into the dialogue and experiences provided.
Detailed character settings and progression frameworks are documented in the Character Profiles section.
We currently need writers and performers who can convincingly inhabit characters. Interested contributors should review the character profiles and writing guidelines.
Programming skills are not required. What matters is lived experience and the ability to reflect real human environments, relationships, and behaviors through writing.
Direct involvement may not be necessary, but individuals with similar backgrounds could potentially help write characters grounded in that experience.
Compensation is currently set at ¥100 per 1,000 words (approximately $14 USD). This may be adjusted based on sustainability and data requirements.