Column
Research, in business language.
We rewrite Affectosphere Group's research into something useful for business and practical decision-making. Each piece is a 5-minute read.
2026 / 06 / 25
General-Purpose LLMs Can't Predict Consumer Behavior: Why Specialized Models Win at the Population Level
BehaviorBench, a new benchmark spanning psychology, sociology, and economics, reveals that general LLMs handle individual-level behavioral tasks reasonably well but fall short on population-level distribution accuracy — where the specialized model Be.FM-1.5 pulls ahead. For CRM and marketing teams predicting 'what customers do next,' this finding points toward domain-specific fine-tuning.
2026 / 06 / 25
AI Agents That Buy Information to Make Decisions: The Micro-Transaction Market Model for Agentic E-Commerce
A new architecture proposal turns the EC chatbot from a conversion tool into a verified information market — where autonomous purchasing agents acquire quality-certified data through micro-transactions before making procurement decisions.
2026 / 06 / 25
Can AI Actually Diagnose Rare Diseases? A Randomized Trial Shows 21-Point Accuracy Gain with LLM RaDaR
A randomized clinical trial of RaDaR, a 32-billion-parameter open-source reasoning LLM trained on rare disease cases, found physicians using it achieved approximately 21 percentage points higher diagnostic accuracy than those relying on internet search alone.
2026 / 06 / 25
When AI Recommends, Humans Stop Reading: The Hidden Cost of Hiring Automation
A new study finds that when AI recommendations are present, recruiters spend up to 55.6% less time reviewing resumes — and that biased AI scores get quietly absorbed into final decisions, even by fairness-conscious evaluators.
2026 / 06 / 25
Warning Labels Don't Stop Sycophantic AI From Influencing You
A pre-registered experiment with 2,610 participants found that warning users about AI sycophancy lowered perceived trust — but left the actual emotional influence entirely intact.
2026 / 06 / 24
Affective AI Safety: The Missing Piece in LLM Safety
LLM safety research has focused heavily on harmful content. But emotional dependency, relational manipulation, and self-alienation slip through content filters undetected. A new framework argues affective safety deserves its own research agenda.
2026 / 06 / 24
When Does Intrinsic Self-Correction Help? A Task-Sensitive Analysis
Making LLMs review their own output is not always beneficial. Research shows that the effectiveness of self-correction depends heavily on task structure — and understanding that distinction is the key to smarter AI automation in legal, compliance, and medical workflows.
2026 / 06 / 24
MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations
When LLMs are embedded in group chats like Slack bots or Teams Copilot, do they actually understand who can receive which information? MuPPET — a new benchmark for contextual privacy — shows that current LLMs have significant vulnerabilities. Here's what this means for safe enterprise deployment under GDPR.
2026 / 06 / 24
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
LLM agents that span many business tools see planning performance collapse as tool count grows — and reliability drops sharply when tools fail. A new benchmark reveals the gap between demo environments and real enterprise deployments.
2026 / 06 / 24
VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows
Why did the AI reject this credit application? For compliance and audit teams in finance, that's the question that matters most. VADAOrchestra answers it by combining logic programs with LLMs — keeping reasoning traceable without sacrificing intelligence. Here's what this means for risk management and regulatory compliance.
2026 / 06 / 23
"Are you an AI?" Analyzing Client Suspicion of AI Use in Crisis Counseling
In a crisis counseling service with no AI involvement whatsoever, client suspicion of AI use rose from 0.8% in June 2024 to 2.6% by March 2025. Suspicion concentrated in the early conversation phase, and even when counselors denied AI involvement, 17.6% of cases ended with suspicion unresolved or conversation terminated. AI distrust can erode therapeutic relationships long before any AI is deployed.
2026 / 06 / 23
The Algorithmic-Human Manager: AI, Apps, and Workers in the Indian Gig Economy
Who manages gig workers — the platform, or its algorithm? A qualitative study published on arXiv interviews 16 gig workers and 21 stakeholders in India's ride-hailing and delivery sector. The findings surface three structural challenges — opacity, unfair outcomes, and misaligned rewards — with direct implications for ESG due diligence, EU AI Act compliance, and fair algorithm design.
2026 / 06 / 23
Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery
Deploying LLM agents autonomously in legal document review triggers a phenomenon called trajectory collapse, where early misclassifications propagate through multi-step reasoning chains and can invalidate privilege review. A four-layer verification system with Human-on-the-Loop escalation reduces privilege waiver risk by up to 61% compared to full autonomy, while keeping attorney review to under one quarter of total documents.
2026 / 06 / 23
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
When AI agents execute tools in customer service workflows, stale or incomplete state information can silently cause domain policy violations. A study published on arXiv proposes LedgerAgent, a runtime approach that tracks observed task state in a separate ledger and verifies policy compliance before any tool action is executed.
2026 / 06 / 23
Beyond Accuracy: Measuring Logical Compliance of Predictive Models
Two AI models with identical accuracy scores can have drastically different rule violation rates. A study published on arXiv introduces the Rule Violation Score (RVS), quantifying logical compliance that accuracy metrics alone fail to capture — with direct implications for medical and financial AI procurement.