logo

Active Learning with Large Language Models for Cross-lingual Pseudo-labeling in Computational Social Science

Location: Austria

Source: EU Funding & Tenders Portal

Project
Ongoing

Large Language Models (LLMs) have become ubiquitous in the area of Natural Language Processing (NLP). This project aims to advance the applications of LLMs in the multidisciplinary field of Computational Social Science (CSS). Specifically, LLMs are often used to pseud-label (at scale) variables in data that are of interest to social scientists (e.g., a psychologist would look for emotion markers

Project Information FAQ

Project Information

3 Q
The project “Active Learning with Large Language Models for Cross-lingual Pseudo-labeling in Computational Social Science” is an infrastructure initiative, located in Austria. Taiyo aggregates data on it from EU Funding & Tenders Portal.

Want to explore the full details? View the full report

Participants

Sponsoring Agency

Obfuscated Data

Company

Obfuscated Data

Status

Original status

ongoing

Taiyo status

Obfuscated Data

Taiyo last update

00-00-0000

Available timestamps

00-00-0000

Available timestamp type

Obfuscated Data

Contact

Contact name

Obfuscated Data

Phone

0000000000

Email

ObfuscatedData@email.com

Address

Obfuscated Data, Obfuscated data, obfuscated data, Obfuscated data

Description

Description

Large Language Models (LLMs) have become ubiquitous in the area of Natural Language Processing (NLP). This project aims to advance the applications of LLMs in the multidisciplinary field of Computational Social Science (CSS). Specifically, LLMs are often used to pseud-label (at scale) variables in data that are of interest to social scientists (e.g., a psychologist would look for emotion markers, an economist for statements about money, a human rights researcher for toxic language etc.) I will develop and test novel tooling for applying LLMs to efficiently pseudo-label CSS data in a cross-lingual setting. The project will proceed in three phases. First, a dataset and methodology for evaluating the language robustness of LLMs (isolated from all other confounding factors) will be developed. Second, using this methodology a set of general best practices will be determined for the scenario of applying LLMs to cross-lingual pseudo-labeling for CSS. Finally, both existing and novel active learning approaches will be investigated to minimize the manual annotation effort required to oversee the pseudo-labeling process. Using insights from this project a user-friendly freely publicly available tool for pseudo-labeling using LLMs will be published. The methods will be tested on two example tasks: toxicity detection and disinformation detection. Overall, the project results will make the work of CSS researchers more efficient in terms of both time and financial resources. Moreover, the cross-lingual nature of the tool will make it applicable to languages other than English, including small and low-resourced languages. Consequently, this project will contribute to make tooling accessible to a wide and diverse set of CSS researchers, increasing the outreach and inclusivity of CSS research and fostering international collaboration.

Original sub-sector

Obfuscated

Original Currency

USD

Original budget

000000000000000

Procurement method

Obfuscated Data

Budget

000000000000000

Location

Region

Obfuscated

Country

Obfuscated

State

Obfuscated Data

County

Obfuscated

Location

Obfuscated Data, Obfuscated data, obfuscated data, Obfuscated data

Source

Source reliability

High

Data quality score

100%

Source

Obfuscated Data

URL

obfuscated_data,obfuscateddata.com

More Details

Project Type

Obfuscated Data

Article Published Date

Obfuscated Data