Appier Research Challenges LLMs to Know When They Don’t Know

Agentic AI Research: When LLMs Know Their Limits Agentic AI Research: When LLMs Know Their Limits

As enterprise AI systems move from answering questions to taking actions, a new problem is becoming harder to ignore: knowing when not to answer. Appier’s AI Research team has published two studies examining whether large language models can recognize missing information and choose the most appropriate language for reasoning—two capabilities that could become important building blocks for more reliable agentic AI.

Appier Research Pushes Agentic AI Beyond Accuracy With Uncertainty and Multilingual Reasoning

The next phase of enterprise AI may depend less on how often a model produces an answer and more on whether it knows when an answer cannot be supported.

That is the central theme behind two research papers published by Appier’s AI Research team, which examine separate but connected problems in large language models (LLMs): recognizing when available information is insufficient and determining which language should be used during reasoning.

The research arrives as enterprises increasingly experiment with AI agents capable of retrieving information, reasoning over it and executing multi-step tasks. Appier, which develops AI technology for advertising and marketing applications, says the findings point toward a broader way of evaluating agentic AI—one that considers uncertainty, reasoning strategy and cultural context rather than answer accuracy alone.

That distinction matters in practical deployments. An AI customer-service agent, for example, could retrieve a return policy for a similar product and incorrectly apply it to a product whose policy is unavailable. A system that simply generates the most plausible answer may appear useful while creating a costly operational problem.

The first Appier study, “None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering,” tested 28 LLMs using questions where “None of the Above” was the correct choice. The researchers found that performance dropped by roughly 30% to 50% when the correct response was that none of the available answers was valid. The decline was particularly pronounced on uncertainty-heavy tasks such as business ethics.

The result highlights a familiar weakness in generative AI: models can be very good at selecting a plausible answer without necessarily establishing that the evidence supports it.

Appier tested two training approaches—Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO)—to improve this behavior. According to the research, DPO increased accuracy in identifying questions without a valid answer by almost 30 percentage points.

For enterprise AI teams, the implication is broader than a benchmark improvement. In a Retrieval-Augmented Generation (RAG) system, retrieval quality is only one part of the reliability equation. An agent also needs a mechanism for deciding whether the retrieved material is sufficient to justify an action.

That could mean introducing a verification checkpoint before an agent responds or executes a workflow. If the evidence is incomplete, the system might retrieve additional information, ask for clarification or hand the case to a human.

Multilingual AI Has Another Reasoning Problem

The second paper explores a different limitation: the language an AI uses internally to reason can affect what it gets right.

In “Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?”, Appier researchers examined multilingual reasoning across mathematical, knowledge, cultural and safety-related tasks. They found that large reasoning models frequently default to high-resource languages such as English even when the input is provided in another language. In some models, the reasoning language and response language differed more than 90% of the time.

The findings are nuanced. Reasoning in a high-resource language generally preserved or improved performance on some mathematical and knowledge tasks. But for cultural understanding, reasoning in the local language could produce better results. The researchers also observed language-specific differences in safety evaluations.

Appier used a technique called text prefilling to influence the language used during reasoning. That points toward a potential architecture the researchers describe as reasoning-language routing: instead of treating language selection as a fixed model setting, an AI agent could choose a reasoning language based on the task, market and cultural context.

For global enterprises, this could become significant. A multilingual customer-support system, international gaming platform or global marketing operation may need to distinguish between language used for interaction and language that produces the strongest reasoning for a particular task.

Why This Matters as AI Agents Move Into Enterprise Workflows

The timing is important. McKinsey’s 2025 State of AI research found that 88% of surveyed organizations were regularly using AI in at least one business function, while 62% said they were at least experimenting with AI agents. Yet most organizations remained in experimentation or pilot phases rather than full-scale deployment.

Gartner has likewise forecast that 40% of enterprise applications could incorporate task-specific AI agents by the end of 2026, up from less than 5% in 2025.

As agents become embedded in business applications, failures that were tolerable in a chatbot become more consequential. An agent that incorrectly answers a casual question is inconvenient. An agent that makes a pricing decision, approves a workflow, interprets a policy or takes action using incomplete evidence can create financial, compliance and reputational risks.

That shifts the competitive conversation around enterprise AI. Model size and benchmark scores still matter, but they are no longer the whole story. Enterprises also need uncertainty detection, retrieval verification, human escalation, multilingual evaluation and governance controls.

Appier plans to continue its LLM and agentic AI research while applying the work across its advertising, personalization and data product lines. The company’s research therefore sits at the intersection of foundational AI capabilities and commercial AI deployment.

The larger takeaway is that reliable agentic AI may require a different definition of intelligence. A useful enterprise agent should not simply produce an answer quickly. It should understand the limits of its evidence, determine how it should reason, and know when a human—or another retrieval cycle—is the better next step.

That is a considerably harder engineering problem than generating fluent text. It may also be one of the more important ones as AI moves from copilots toward autonomous enterprise systems.

Market Landscape

Enterprise AI is shifting from isolated copilots toward systems that can retrieve information, reason across multiple steps and take action. McKinsey reports that 88% of surveyed organizations were using AI in at least one business function in 2025, but most companies had not yet scaled AI across the enterprise.

At the same time, Gartner expects task-specific AI agents to become significantly more common inside enterprise applications.

That creates an emerging differentiation point for AI platforms. Microsoft, Google, Amazon, Salesforce and other enterprise technology vendors are building increasingly agentic systems, while model developers continue competing on reasoning, tool use and multimodal capabilities. The next layer of competition is likely to involve how reliably these systems operate under uncertainty.

Appier’s research adds two dimensions to that discussion: whether an AI agent can decline unsupported answers and whether it can adapt its reasoning process to multilingual and culturally specific tasks.

For enterprise buyers, the practical question is therefore moving from “Which model is smartest?” toward “Which system can fail safely when the model does not know?”

Top Insights

  • Appier’s LLM research found major performance declines when “None of the Above” was correct, highlighting uncertainty recognition as an enterprise AI reliability challenge.
  • DPO training improved detection of questions without valid answers, suggesting targeted alignment can strengthen an agent’s ability to withhold unsupported decisions.
  • Multilingual reasoning research found that language choice affects mathematical, cultural and safety performance, complicating global deployment of large reasoning models.
  • Reasoning-language routing could allow AI agents to dynamically select reasoning languages based on task requirements, markets and cultural context.
  • As enterprise agents move toward autonomous workflows, evidence verification, uncertainty handling and human escalation may become as important as model accuracy.

Power Tomorrow’s Intelligence — Build It with TechEdgeAI

Grow Your
Brand Visibility

Looking to publish a press release, guest article, interview or podcast? Connect with us.

GET FEATURED
Subscribe

Sign up today for exclusive insights and updates.

Newsletter Signup