Multilingual Data Contributors: PDF Collection for AI Training
$150/task
source wording: “Terac listing record, payRate 150.00 USD PER_TASK”
What this role asks for
- Fluent Gujarati
- Fluent Malayalam
- Fluent Telugu
- Data Collection
- AI Training
- Multilingual Research
- Document Sourcing
- Machine learning
Read from the posting's own words: show the exact sentences
- Fluent Gujarati: “…guidelines before approving the task. Who This Is For We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents online. Ideal…”
- Fluent Malayalam: “…guidelines before approving the task. Who This Is For We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents online. Ideal candidates…”
- Fluent Telugu: “…guidelines before approving the task. Who This Is For We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents…”
- Data Collection: “listed by the platform under required skills: Data Collection”
- AI Training: “listed by the platform under required skills: AI Training”
- Multilingual Research: “listed by the platform under required skills: Multilingual Research”
- Document Sourcing: “listed by the platform under required skills: Document Sourcing”
- Machine learning: “…usable PDFs in various languages are essential for training robust machine learning models. Your contributions will directly support the development of better language…”
The role, as Terac describes it
We are running a paid project to collect legally usable PDF documents to help train artificial intelligence models. We are looking for fluent speakers of specific Asian languages to source and submit high-quality text files.
What We're Researching
We're running a paid study on multilingual document sourcing to improve AI text recognition and generation. High-quality, legally usable PDFs in various languages are essential for training robust machine learning models. Your contributions will directly support the development of better language processing tools. How It Works
Read the rest of the description (12 more paragraphs)
You will work asynchronously to find and submit public, legally usable PDF documents in your designated language. During this process, you will verify that each document meets our quality and licensing requirements. You will upload the files through our secure platform and provide basic metadata for each submission. We will review your uploaded documents to ensure they match the project guidelines before approving the task. Who This Is For
We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents online. Ideal candidates are detail-oriented individuals comfortable navigating digital archives, public records, or open-source repositories. We welcome data annotators, researchers, and general language contributors who understand basic copyright and licensing rules.
What you would do
• Source legally usable, public PDF documents in your designated language
• Verify that each document meets open-source or public domain licensing requirements
• Upload the collected files to our research platform
• Provide basic descriptive information for each submitted document
Who this is for
• Fluent reading comprehension in Telugu, Odia, Gujarati, Malayalam, Japanese, or Korean
• Comfortable searching for and downloading digital documents online
• Basic understanding of public domain or open-source licensing
• Access to a reliable computer and internet connection for uploading files
Posted by Terac, reproduced here so you can judge the role before clicking. Original posting ↗
Source: platform job feed · first seen 4d ago · ID x9sLPDNsxxpmTjO5
More like this
- Apply ↗AI Facial Data Collection Contributorup to $140/hrmicro1Data annotationWorldwideposted 25d ago
$150/taskTerac · Multilingual Data Contributors: PDF Collection for AI Training
Apply on Terac ↗