E

Data Scientist (AI Data & LLM Specialist)

Eclipse · 公链/L2 · 上架 2026-05-28
面议
数据/AI中级📍 Remote远程

职位要求 / 描述

<p>Join the core team at Eclipse, where we’re building an AI agent-first marketplace that connects intelligence with real-world tasks, starting with data collection and labeling. We are seeking a Data Scientist to establish the foundation for how our data is labeled, processed, and prepared for consumption by next-generation Large Language Models (LLMs). Your work will be critical in transforming our raw data collections into valuable, AI-ready datasets.</p> <h2><strong>Qualifications </strong></h2> <ul> <li>Proven experience as a Data Scientist or Machine Learning Engineer with a focus on data quality and preparation.</li> <li>Strong understanding of data labeling methodologies and hands-on experience with data annotation platforms and workflows.</li> <li>Demonstrated experience preparing datasets for training and fine-tuning Large Language Models (LLMs), including knowledge of techniques like tokenization, embeddings, and NER.</li> <li>Proficiency in Python and common data science libraries (e.g., Pandas, NumPy, Scikit-learn, spaCy, Hugging Face).</li> <li>Experience using APIs/SDKs to automate data annotation and active learning loops.</li> <li>Excellent communication skills, with an ability to create clear documentation for technical and non-technical audiences.</li> </ul> <h2><strong>Responsibilities </strong></h2> <ul> <li>Develop Data Labeling Strategies: Design and document a formal data annotation strategy, including clear, scalable, and efficient guidelines for labeling our data. Define and enforce quality metrics, including inter-annotator agreement.</li> <li>Optimize for LLM Consumption: Research, define, and prototype the optimal data formats, structures, and pre-processing steps required for fine-tuning and training LLMs on our datasets.</li> <li>Data Quality Analysis: Establish automated processes and metrics to analyze the quality of both raw and labeled data, providing feedback to improve our data collection and labeling workflows.</li> <li>Collaborate with Engineering: Work closely with the engineering team to guide the implementation of data processing pipelines and ensure the data infrastructure meets the needs of ML applications.</li> </ul> <h2><strong>Nice-to-Haves</strong></h2> <ul> <li>Experience with audio data processing and relevant libraries.</li> <li>Familiarity with data annotation platforms and tools.</li> <li>Knowledge of modern MLOps principles and practices.</li> <li>Experience with large language model data curation and Reinforcement Learning from Human Feedback (RLHF) pipelines.</li> </ul> <h2><strong>Join the Eclipse team!</strong></h2> <p>Eclipse is building the fastest Ethereum Layer 2, powered by the Solana VM. Our general-purpose L2 combines the best of the modular stack without sacrificing UX or fragmenting liquidity. On top of this foundation, we’re building apps in-house and iterating quickly to find breakout consumer and AI experiences. We’re backed by top investors including Polychain, Tribe Capital, Placeholder, and DBA.</p> <ul> <li><strong>Opportunity</strong>. We believe blockchains should be fast AND highly usable. You’ll do high-impact work to enhance Ethereum’s scalability, shaping the future of crypto</li> <li><strong>Flexibility</strong>. We collaborate synchronously and asynchronously, across weekly all-hands meetings, Slack messaging, and quarterly in-person meetups</li> <li><strong>Team</strong>. Our founding team has experience launching and scaling blue-chip projects such as dYdX, Uniswap, and zkSync. We’re backed by leading funds and leaders including Polychain, Tribe, Placeholder, DBA, Mustafa Al-Bassam, Tarun Chitra, Meltem Demirors, and others</li> <li><strong>Culture</strong>. As an early member of our team, you’ll have a unique opportunity to help shape our culture. We value intellectual honesty, bias towards action, and believe every member plays a key role in achieving our ambitious goals</li> <li><strong>Compensation</strong>. You’ll receive a competitive sal

技能关键字

#Python#L2/Rollup#Ethereum#Solana#Machine Learning#AI

职责方向

性能/容量部署发布自动化数据分析设计/品牌

数据来自公开渠道整理,薪资为公开 JD 或聚合估算,仅供参考,以面试谈薪为准。 ← 返回链聘 ChainHire 职位看板

Data Scientist (AI Data & LLM Specialist) · Eclipse
面议
立即投递 →