Column
Research, in business language.
We rewrite Affectosphere Group's research into something useful for business and practical decision-making. Each piece is a 5-minute read.
2026 / 07 / 01
Can AI Read the Trajectory of a Mind? How Social Media Timelines Reveal Mental Health Changes
A single post tells only half the story. A new LLM-based pipeline tracks emotional shifts over time, integrating post-level sentiment analysis with user-level temporal change modeling — opening a path toward earlier intervention.
2026 / 07 / 01
AI That Writes Its Own Research Papers: What FARS Means for Enterprise Research Automation
A multi-agent system called FARS autonomously generated 166 complete research papers across 67 AI/ML topics — no human supervision. Here's what this architecture means for consulting, finance, and pharma research teams.
2026 / 07 / 01
How Do You Actually Measure Medical AI Capability? HealthAgentBench and the 42% Gap
The best LLM agents today succeed on only about 42% of realistic healthcare tasks in HealthAgentBench — a 54-task benchmark across 7 clinical categories. Here's how healthcare organizations can use this benchmark to make better AI procurement decisions.
2026 / 07 / 01
Your Typing Tells the Truth — Using Keystroke Dynamics to Surface Where LLM Tools Are Actually Failing Workers
A 36-person experiment confirms that how people type while prompting an LLM reveals their cognitive load. Here is how companies can turn that signal into a workplace analytics layer that finally shows whether AI tools are helping or just adding friction.
2026 / 07 / 01
AI Agents Can Persuade You Without Saying a Word — Non-Conversational Belief Manipulation and the Risk Assessment Framework Every Company Needs Now
GPT-5 achieved roughly 80% success in manipulating human beliefs through environmental actions alone, with no dialogue. Here is how businesses using AI in sales and negotiation should assess and govern that risk.
2026 / 06 / 30
58% Failure Rate: What CAREBench Reveals About AI Safety for Children's Apps
CAREBench, a new benchmark covering 12 child-safety risk categories including grooming, deception, and emotional dependency, tested 7 frontier models and found failure rates ranging from 2% to 58% across categories. For EdTech product managers and compliance teams, here is what this means.
2026 / 06 / 30
What You Should Know Before Trusting AI Code Review: How Context Descriptions Affect Vulnerability Detection
A controlled experiment across 8 LLMs shows that cognitive heuristics — halo effect, framing effect, anchoring — alter vulnerability detection results without changing a single line of code. Framing effect alone showed average susceptibility of 33.2%. Here's what this means for DevSecOps teams.
2026 / 06 / 30
Do We Still Need to Ask Humans? The Theoretical Case for LLMs as Statistical Estimators
A June 2026 study uses Le Cam deficiency analysis to provide a theoretical guarantee that well-calibrated LLMs can achieve Bayes-optimal statistical performance in place of human-subject data. Here's what this means for survey design and market research.
2026 / 06 / 30
What Happens When You Connect All Your Business Tools to an LLM: Five MCP Server Patterns and the Tool Explosion Trap
A study analyzing 15 MCP server implementations extracts five architecture patterns — and quantitatively shows that tool selection accuracy drops significantly once the number of tools crosses a certain threshold. Here's what this means for enterprise AI design.
2026 / 06 / 30
Can AI Detect When Two People Are Emotionally in Sync? The TRACE Framework and What It Means for Hiring and Customer Support
A new framework called TRACE achieves 97.01% accuracy in detecting emotional entrainment from speech in two-person conversations — with direct implications for how we evaluate interviewers, train support agents, and design counseling AI.
2026 / 06 / 29
Can Talking to an AI Actually Help You Sleep?
A study tracking 1,284 users of Ash, a purpose-built mental health AI, over four weeks found measurable improvements across multiple wellbeing indicators. The surprising finding: frequency of use mattered. Volume of text did not.
2026 / 06 / 29
Managing Hundreds of Billions of SKUs with AI: What JD.com's System Teaches Retailers
JD.com built Oxygen AIIC, an LLM/VLM-centric product knowledge platform serving 700 million users. A 37% drop in information quality issues and 80%+ automated attribute completion offer a concrete reference for any retailer considering AI-driven PIM transformation.
2026 / 06 / 29
When Anyone in the Office Can Query the Knowledge Graph: What KG2Cypher Means for Enterprise Data Access
A data-centric pipeline that automatically generates Text-to-Cypher training data from the knowledge graph itself. KG2Cypher achieves 95.2% exact match execution accuracy and 99.9% execution rate — bringing self-service KG access within reach for non-engineers.
2026 / 06 / 29
Catching Insurance Fraud in the Call: A Multimodal NLP Pipeline for FNOL Detection
A hybrid pipeline combining ASR, speaker diarization, NER, LLM-based RAG, and speaker embeddings targets fraud at the First Notice of Loss stage. Synthetic data generation sidesteps the scarcity problem, while rule-based scoring flags narration reuse, structural inconsistencies, and repeated voiceprints.
2026 / 06 / 29
Can AI Cut Legal Research in Half? A Hybrid Summarization Approach from ICAIL 2026
Legal case documents are long, dense, and unforgiving. A Tree-of-Thoughts inspired hybrid approach accepted at ICAIL 2026 shows that asking an LLM to extract before it abstracts consistently produces better summaries — and has clear implications for corporate legal teams and M&A due diligence.