| Job Position and Company | Posted | Location | Salary | Tags |
|---|---|---|---|---|
|
| ||||
|
| ||||
Bcbgroup📍 Remote | $105k - $120k | |||
Zinnia📍 Remote | $91k - $99k | |||
ISO 9001 Certified | 400+ students | Learn more | by Metana | ||
Okx📍 Remote | $98k - $150k | |||
Integra📍 Remote | $21k - $52k | |||
Integra📍 Remote | $21k - $52k | |||
| $72k - $84k | ||||
| $88k - $101k | ||||
| $105k - $108k | ||||
Bluecubeservices📍 Remote | $79k - $100k | |||
Alpaca📍 Remote | $122k - $123k | |||
Nansen📍 Remote | $72k - $100k | |||
| $88k - $101k | ||||
| $400k - $500k |
Financial Data Engineer, AI/LLM
About Binance
Binance is the global leading blockchain ecosystem, operating the world's largest digital asset trading platform by volume, serving over 300 million users across 100+ countries and regions. We are committed to building a more open financial ecosystem and improving global access to financial services.
Binance is continuously building stock and related financial market products for global users. The relevant data will serve user-facing stock products and Binance AI business scenarios. We are seeking professionals with stock market experience to jointly build reliable, scalable data and financial AI capabilities.
Role Overview
You will participate in building the core data foundation for Binance's stock and related financial market businesses, responsible for the full pipeline from data source discovery, evaluation, ingestion, and integration to unified modeling, real-time processing, quality governance, and data services. Beyond completing defined integrations, we expect you to leverage industry expertise to continuously identify better data sources and technical solutions, enabling new markets, products, and data to serve trading products and AI quickly and reliably.
Responsibilities
- Conduct research, technical evaluation, ingestion, integration, cleansing, standardization, computation, storage, and servicing of financial market data, covering securities master data, real-time and historical market data, fundamentals, corporate actions, indices, and business-required product and risk data; responsible for source ingestion, raw retention, and stable delivery to knowledge engineering pipelines for content-type data such as announcements, news, and research reports.
- Design scalable unified data models and integration frameworks, handling different markets' trading calendars, time zones, currencies, security identifiers, listing relationships, lifecycles, and data corrections, supporting rapid onboarding of new markets and sources.
- Build batch-stream unified data pipelines centered on Flink, continuously optimizing latency, throughput, query performance, stability, and cost, while supporting consumer trading products, research analysis, and AI scenarios.
- Establish data quality and service level frameworks, taking responsibility for completeness, accuracy, timeliness, consistency, and traceability; build automated reconciliation, anomaly detection, monitoring and alerting, raw data replay, backfill, and fault recovery capabilities.
- Evaluate different sources (vendors, exchanges, APIs, file feeds, compliance collection) for coverage, quality, stability, revision mechanisms, and technical fit; collaborate with product, procurement, legal, and compliance teams to clarify usage, display, derivative, retention, and redistribution boundaries; drive rational primary/backup source strategies and alternatives.
- Define data semantics, metric definitions, and service contracts jointly with trading product, data platform, AI engineering, and algorithm teams, ensuring consistent and reliable usage of the same stock facts across different products.
- Drive data engineering efficiency and technical quality improvements, including metadata, data lineage, automated testing, CI/CD, task orchestration, capacity governance, and AI-assisted development.
Requirements
- Master's degree or above in Computer Science, Software Engineering, Mathematics, Statistics, or related field; 5+ years of experience in data development, big data, or data platforms.
- Familiar with stock markets and the investor research and decision-making workflow; understand trading mechanisms, market data, fundamentals and financial reports, corporate actions, valuation, and major market events; able to explain the complete pipeline of at least one type of financial data from source to user-facing product and key quality risks.
- Proficient in SQL and Flink, with experience in large-scale real-time data processing, performance tuning, stability governance, and production issue troubleshooting.
- Proficient in at least one of Java, Scala, or Python; familiar with Kafka, Spark, and ClickHouse, Doris, HBase, Elasticsearch, or other distributed storage and analytics technologies.
- Familiar with data modeling, task scheduling, metadata, data lineage, data governance, and service levels; able to independently resolve cross-system data consistency issues.
- High standards for data quality; able to design reproducible reconciliation, anomaly detection, backfill, and degradation strategies — not just completing data development tasks.
- Experience with data source selection or production ingestion; able to articulate trade-offs between build vs. buy, multi-source verification, vendor dependency, and alternative solutions.
- Strong business understanding and cross-team collaboration skills; able to translate trading, risk, research, or AI problems into clear data models and data contracts.
Bonus
- Experience with stock data at brokerages, market data services, financial data, wealth management, or fintech platforms.
- Familiarity with US stock market structure, trading calendars, pre/post-market sessions, corporate actions, and adjustment rules; experience with other stock markets also a plus.
- Data experience with stock-related derivatives, ETFs, indices, or tokenized products.
- Experience building low-latency market data pipelines, securities master data platforms, multi-market data models, quantitative research platforms, or large-scale backtesting data systems.
- Experience with data anomaly detection, knowledge graphs, financial entity alignment, or building high-quality financial datasets for LLMs and retrieval-augmented generation (RAG).
What does a data scientist in web3 do?
A data scientist in web3 is a type of data scientist who focuses on working with data related to the development of web-based technologies and applications that are part of the larger web3 ecosystem
This can include working with data from decentralized applications (DApps), blockchain networks, and other types of distributed and decentralized systems
In general, a data scientist in web3 is responsible for using data analysis and machine learning techniques to help organizations and individuals understand, interpret, and make decisions based on the data generated by these systems
Some specific tasks that a data scientist in web3 might be involved in include developing predictive models, conducting research, and creating data visualizations.