logo

Non-sequence models for tokenization replacement

Sector: Education • Location: Germany

Source: EU Funding & Tenders Portal

Project
Ended

Natural language processing (NLP) is concerned with computer-based processing of natural language, with applications such as human-machine interfaces and information access. The capabilities of NLP are currently severely limited compared to humans. NLP has high error rates for languages that differ from English (e.g., languages with higher morphological complexity like Czech) and for text genres

Project Information FAQ

Project Information

3 Q
The project “Non-sequence models for tokenization replacement” is an infrastructure initiative in the Education sector, located in Germany. Taiyo aggregates data on it from EU Funding & Tenders Portal.

Want to explore the full details? View the full report

Participants

Sponsoring Agency

Obfuscated Data

Company

Obfuscated Data

Status

Original status

ended

Taiyo status

Obfuscated Data

Taiyo last update

00-00-0000

Available timestamps

00-00-0000

Available timestamp type

Obfuscated Data

Contact

Contact name

Obfuscated Data

Phone

0000000000

Email

ObfuscatedData@email.com

Address

Obfuscated Data, Obfuscated data, obfuscated data, Obfuscated data

Description

Description

Natural language processing (NLP) is concerned with computer-based processing of natural language, with applications such as human-machine interfaces and information access. The capabilities of NLP are currently severely limited compared to humans. NLP has high error rates for languages that differ from English (e.g., languages with higher morphological complexity like Czech) and for text genres that are not well edited (or noisy) and that are of high economic importance, e.g., social media text. NLP is based on machine learning, which requires as basis a representation that reflects the underlying structure of the domain, in this case the structure of language. But representations currently used are symbol-based: text is broken into surface forms by sequence models that implement tokenization heuristics and treat each surface form as a symbol or represent it as an embedding (a vector representation) of that symbol. These heuristics are arbitrary and error-prone, especially for non-English and noisy text, resulting in poor performance. Advances in deep learning now make it possible to take the embedding idea and liberate it from the limitations of symbolic tokenization. I have the interdisciplinary expertise in computational linguistics, computer science and deep learning required for this project and am thus in the unique position to design a radically new robust and powerful non-symbolic text representation that captures all aspects of form and meaning that NLP needs for successful processing. By creating a text representation for NLP that is not impeded by the limitations of symbol-based tokenization, the foundations are laid to take NLP applications like human-machine interaction, human-human communication supported by machine translation and information access to the next level.

Original sub-sector

Obfuscated

Original Currency

USD

Original budget

000000000000000

Procurement method

Obfuscated Data

Budget

000000000000000

Location

Region

Obfuscated

Country

Obfuscated

State

Obfuscated Data

County

Obfuscated

Location

Obfuscated Data, Obfuscated data, obfuscated data, Obfuscated data

Source

Source reliability

High

Data quality score

100%

Source

Obfuscated Data

URL

obfuscated_data,obfuscateddata.com

More Details

Project Type

Obfuscated Data

Article Published Date

Obfuscated Data