hirq
← All jobs

GT

Senior Applied Data Scientist | NDA

Warsaw, Poland · Remote · FullTime · Confidential

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

Machine LearningLLMsA/B TestingPythonSQLSnowflakeBigQueryDatabricksSparkNLPPyTorchTensorFlow
GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands. On behalf of our client, GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML, embedding, and LLM-based approaches to solve complex data matching problems at scale. ABOUT THE CLIENT Our client is a leading global management consultancy known for tackling some of the world’s most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact. ABOUT THE ROLE We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale. You will develop and test new ML, embedding, and LLM-based approaches for matching complex business records across multiple data sources. The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches. A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets. RESPONSIBILITIES: Develop better ways to match company records - Build new ML, embedding, and LLM-based approaches for matching entities - Improve how the system handles messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies. - Develop scoring and ranking approaches to distinguish accurate matches from duplicates, similar-looking records, and unrelated entities. - Evaluate and implement AI and machine learning techniques to improve matching quality while considering accuracy, scalability, and cost. - Design approaches that can operate efficiently at scale, taking model usage and computational cost into consideration. Improve evaluation, experimentation, and match quality - Define and improve methods for evaluating match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review effort. - Assist in building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout. - Explore LLM-assisted review and validation to assess matching performance and benchmark more scalable approaches. - Turn ambiguous matching problems into clear hypotheses, experiments, metrics, and recommendations. Partner with engineering to bring successful ideas into production - Work closely with data engineering and software engineering teams to turn promising prototypes into production-ready matching logic. - Provide engineering partners with clear model specifications, evaluation results, expected behavior, edge cases, and rollout requirements. - Help determine the most appropriate matching techniques based on data characteristics, confidence levels, and cost considerations. - Continuously evaluate matching performance, investigate regressions, and recommend improvements to models and matching logic. - Clearly communicate technical tradeoffs related to matching performance, scalability, cost, latency, explainability, and operational considerations. ESSENTIAL KNOWLEDGE, SKILLS & EXPERIENCE: - 5–8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role. - Strong applied ML fundamentals, with hands-on experience building and evaluating models on real data. - Excellent Python and SQL skills. - Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques. - Hands-on experience training supervised and unsupervised models, including classification and NLP tasks. - Working knowledge of neural network and transformer architectures. - Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret. - Experience retraining a taxonomy classifier or maintaining classification models in production. - Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved. - Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners. NICE-TO-HAVE: - Experience with entity resolution, record linkage, deduplication, or similar matching problems. - Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation. - Experience applying LLMs or embeddings to business problems where cost and scale matter. - Exposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQuery. - Familiarity with company, domain, website, firmographic, or other business-entity data. INTERVIEW STEPS: 1. GT interview with Recruiter 2. Technical interview 3. Final interview