Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

AI / ML

AI Data Engineer

Flam

FresherOn-site · Bengaluru, Karnataka, IndiaFull-timeListed 20d ago
Apply on LinkedIn

How to stand out for AI Data Engineer at Flam

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience Fresher-friendly - open to freshers / early-career candidates

About the role

structured by ORI

Flam is building the next generation of interactive media through its content format. We are an AI-native technology company transforming how brands and consumers interact through immersive, interactive content.

What you will do

  • Build ingestion and cleaning pipelines at scale — deduplication (exact, near-dup, semantic), quality filtering with classifier-based scoring, PII stripping,language identification, format normalization.
  • Generate synthetic data where real data is scarce or expensive: LLM generated instruction and preference data, code-mixed Indic text, rendered scenes for image models, augmented audio for speech models.
  • Curate multilingual and code-mixed Indic corpora.
  • Build and maintain evaluation sets. Design them so they measure what we think they measure, keep them uncontaminated, and version them properly
  • Handle multimodal data - audio segmentation and transcript alignment for TTS/ASR, face and video preprocessing for avatar training, image-caption pair curation.

What they are looking for

  • Strong Python. You are comfortable processing datasets far larger than memory, and you know when to reach for a database instead of a script.
  • Practical data engineering fundamentals — streaming, chunking, parallelism, checkpointing long jobs, handling malformed input without losing the run.
  • Familiarity with the modern data formats and tooling of ML Parquet,WebDataset, HuggingFace Datasets, object storage S3/GCS/R2.
  • Genuine care about data quality. You should find it uncomfortable when a dataset has duplicates in it.
  • Enough ML understanding to know how a data decision propagates into model behaviour — why dedup matters, why eval contamination invalidates a benchmark, what a bad filter does to a distribution's tails.

Nice to have

  • Experience curating training data for LLMs, speech models, or image/video models specifically.
  • Working knowledge of embeddings and vector search for semantic deduplication and retrieval FAISS, Milvus, or similar).
  • Native or near-native fluency in one or more Indian languages, with the ability to judge quality in it — this is a real asset here, not a checkbox.
  • Audio or video processing experience (ffmpeg, torchaudio, forced alignment).
  • Synthetic data generation using LLMs, or 3D rendering pipelines Blender, Unreal) for visual data.

Benefits

  • Support to move into model training over time
PythonDeduplicationPII StrippingLanguage IdentificationSynthetic Data GenerationAudio SegmentationTranscript AlignmentVector SearchParquetWebDatasetHuggingFace DatasetsS3GCSR2FAISS
Full posting text

Flam is building the next generation of interactive media through its content format. We are an AI-native technology company transforming how brands and consumers interact through immersive, interactive content. Our technology enables rich, app-less experiences that can be launched instantly on smartphones, creating a fundamentally different way for brands to engage consumers. We are backed by leading investors and already work with some of the world's largest brands. We are now building Flicks, our interactive media format for the US market.

About the role :

Every model we ship at Flam is bounded by the data that trained it, and every eval number we report is only as trustworthy as the eval set behind it. This role owns both. You will build and run the data layer underneath our LLM Falcon, our TTS system Finesse). That means text, audio, and video, cleaned, deduplicated,filtered,labelled, synthetically generated where real data doesn't exist, and versioned so that six months from now we can say exactly what went into a checkpoint.

This is an engineering role. You will write Python every day. It is not an annotation or labelling-management position, though you will design annotation guidelines and quality-check what comes back.

What you'll do

Build ingestion and cleaning pipelines at scale — deduplication (exact, near-dup, semantic), quality filtering with classifier-based scoring, PII stripping,language identification, format normalization.

Generate synthetic data where real data is scarce or expensive: LLM generated instruction and preference data, code-mixed Indic text, rendered scenes for image models, augmented audio for speech models.

Curate multilingual and code-mixed Indic corpora. A large part of our differentiation is Indic-language quality, and a large part of that comes down to data hygiene most pipelines get wrong.

Build and maintain evaluation sets. Design them so they measure what we think they measure, keep them uncontaminated, and version them properly

Handle multimodal data - audio segmentation and transcript alignment for TTS/ASR, face and video preprocessing for avatar training, image-caption pair curation.

Own data provenance and versioning. Which files, which filters, which version, which run. This should be answerable in one command, not one

afternoon. Write annotation guidelines and audit annotation quality when human labelling is in the loop.

What we're looking for:

Strong Python. You are comfortable processing datasets far larger than memory, and you know when to reach for a database instead of a script.

Practical data engineering fundamentals — streaming, chunking, parallelism, checkpointing long jobs, handling malformed input without losing the run.

Familiarity with the modern data formats and tooling of ML Parquet,WebDataset, HuggingFace Datasets, object storage S3/GCS/R2.

Genuine care about data quality. You should find it uncomfortable when a dataset has duplicates in it.

Enough ML understanding to know how a data decision propagates into model behaviour — why dedup matters, why eval contamination invalidates a benchmark, what a bad filter does to a distribution's tails.

Strongly preferred:

Experience curating training data for LLMs, speech models, or image/video models specifically.

Working knowledge of embeddings and vector search for semantic deduplication and retrieval FAISS, Milvus, or similar).

Native or near-native fluency in one or more Indian languages, with the ability to judge quality in it — this is a real asset here, not a checkbox.

Audio or video processing experience (ffmpeg, torchaudio, forced alignment).

Synthetic data generation using LLMs, or 3D rendering pipelines Blender, Unreal) for visual data.

Not required:

A degree in ML.

Prior model training experience.

Experience with our exact toolchain.

Why this role matters :

At most companies, data work is what gets handed to whoever is available. Here it's a named role with a named owner, because our model quality is directly downstream of it. You'll work alongside the engineers training the models, and you'll see your filtering decisions show up in benchmark numbers within the same quarter. If you want to move into model training over time, this is a strong path into it and we'll support that.

Seniority level: Entry level

Employment type: Full-time

Job function: Information Technology

Industries: Technology, Information and Media

Information TechnologyTechnology, Information and Media
Apply on LinkedIn

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…

More at Flam

Jobgether

On-site · India

3+ yrs · 1h ago

Platform Engineer & Cloud Ops Engineer
View role
Flexiple

Remote · India

3–7 yrs · 1h ago

DevOps Engineer
View role
Vantive

On-site · Bengaluru, Karnataka, India

Senior · 10+ yrs · 1h ago

JDE DevOps Consultant
View role
InfosysPreferred

On-site · Hyderabad, Telangana, India

5–8 yrs · 1h ago

AWS Terrafrom DevOps
View role
Apply on LinkedIn