We use essential cookies for site functionality. With your consent, we also use analytics cookies (Google Analytics) to improve Wise Racer. You can change your choice anytime via the Manage Cookies link in the footer. See our Privacy & Cookie Policy

Wise Racer
HomeBlogContact Us

Are Swimming’s Fitness and Competitive Industries Data-Fit for AI? – Part 1

Are Swimming’s Fitness and Competitive Industries Data-Fit for AI? – Part 1

Published on February 11, 2025
Edited on July 7, 2026


Introduction

Data-driven insights have changed many sports, supporting more precise training plans, injury-risk monitoring, and real-time performance feedback (Vec et al., 2024; Leckey et al., 2025). Yet, in the realm of swimming—a sport where milliseconds matter—the quality and structure of data remain significant challenges. How can AI and ML help us unlock better decisions, and what risks arise when data quality is ignored?

This first installment of our two-part series offers a literature-based review on preparing data for AI in sports. It draws from AI/ML research fields and applies those ideas to swimming-specific scenarios. Our goal is to bridge the gap between what AI systems need and how swimming can provide it. We will explore the foundations of data quality, the dangers of poor data management, and the key pillars needed to build robust, AI-ready datasets. By the end of this review, you will understand why well-structured, high-quality data is essential for advanced analytics, better decision-making, and more useful performance feedback in the pool.

Sections Covered in Part 1:

  • Section 1: Why Data Quality Is Essential for ML/AI We outline the core reasons why high-quality, well-managed data is indispensable for AI and ML applications, especially in performance-critical sports like swimming.
  • Section 2: The Barriers, Pitfalls, and Challenges of Poor-Quality Data This section highlights the practical consequences of poor data practices, including biased models, flawed training strategies, and wasted resources.
  • Section 3: Core Foundations for Ensuring High-Quality Data in AI/ML We present the key pillars of reliable data management, from intrinsic and contextual data quality to ethical compliance, all of which are crucial for creating trustworthy AI outcomes.

Section 1: Why Data Quality Is Essential for ML/AI — “The Engine of AI”

Imagine an engine running on low-grade or contaminated fuel. It cannot perform at its best. Data works in a similar way for Machine Learning (ML) and Artificial Intelligence (AI). In sport, especially swimming, accurate data powers modern analytics, performance tracking, and decision-making. Poor-quality or incomplete data can mislead even the most advanced AI systems, potentially distorting training plans and competitive decisions.

Below are key reasons why data quality is vital for any AI-driven application:

  1. Model Accuracy and Reliability High-quality data helps AI models produce more reliable outputs. In swimming, consistent and accurate data on metrics like stroke counts, lap splits, and heart-rate variability can help coaches and athletes interpret AI-generated insights with more confidence. On the other hand, poor data can lead to unreliable models and flawed training recommendations (Priestley et al., 2023; Polyzotis et al., 2018).
  2. Avoidance of Data Cascades Data errors can propagate throughout the ML pipeline, creating a cascade effect where small initial mistakes amplify into larger problems. For instance, consistently misrecording lap times may distort pace analysis, fatigue predictions, and race strategies, leading to costly inefficiencies (Sambasivan et al., 2021; Polyzotis et al., 2018).
  3. Bias and Fairness Biased or incomplete data, especially in competitive sports, can result in skewed insights and inequitable outcomes. For example, training data limited to certain swimmer demographics or conditions may exclude key factors, creating models that favour some athletes over others. Ensuring diverse, representative data helps reduce bias and improve generalization (Zhou et al., 2024; Qayyum et al., 2020).
  4. Data Cleaning and Preparation Effective data cleaning removes noise, corrects inconsistencies, and addresses missing values. Think of it as maintaining a pool's water quality—without proper cleaning, swimmers' performance data and AI insights can suffer. Cleaning for ML should also consider the intended application, not just generic error removal (Neutatz et al., 2021; Polyzotis et al., 2018; Priestley et al., 2023).
  5. Domain-Specific Requirements Each sport comes with unique metrics and requirements. In swimming, monitoring metrics like stroke frequency, rest intervals, and underwater phases is essential. Tailoring data quality checks to these specifics helps AI outputs address real-world performance needs rather than generic data availability alone (Priestley et al., 2023; Rangineni, 2023).
  6. Continuous Monitoring and Management Data collection does not stop after a model is trained. Swimmers' performance evolves, new athletes join programs, and sensors may change over time. Ongoing monitoring of incoming data helps AI tools remain relevant to the context in which they are being used (Bangad et al., 2024; Zhou et al., 2024).
  7. Comprehensive Data Quality Management Managing large volumes and varieties of training data—such as lap counts, biometric readings, and video analytics—requires robust, scalable processes. A clear data quality strategy addresses volume, variety, and velocity to maintain consistency across the ML lifecycle (Rangineni, 2023; Priestley et al., 2023).
  8. Ethical and Legal Considerations Collecting performance and health metrics raises ethical concerns, especially around privacy and compliance. High data quality standards, secure management, and adherence to ethical guidelines help organizations meet legal obligations (Qayyum et al., 2020; Zhou et al., 2024).

Data quality is the foundation of successful ML/AI systems. Accurate, comprehensive, and well-managed data can drive more reliable models, fostering trust among coaches, athletes, and stakeholders. Treating data as the “fuel” of AI applications supports more equitable outcomes, whether in training facilities, research labs, or global competitions.

Section 2: The Barriers, Pitfalls, and Challenges of Poor-Quality Data

In sports analytics, poor data quality is more than just a minor setback—it can derail training programs, waste valuable resources, and erode trust in AI-driven insights. From coaches tracking turn times to sports scientists analyzing large sensor datasets, understanding these key pitfalls is crucial to ensuring reliable outcomes.

  1. Model Performance Degradation AI models rely on accurate, complete data to learn and make predictions. When fed missing or incorrect data—such as inaccurate lap splits or mislogged stroke counts—models can produce unreliable predictions. This can result in suboptimal pacing strategies or potentially inappropriate training decisions if outputs are treated as prescriptions rather than coach-reviewed decision support (Priestley et al., 2023; Leckey et al., 2025).
  2. Data Cascades Small data errors at the start of the pipeline can snowball into larger issues downstream. For example, a heart-rate monitor that incorrectly records frequent spikes could trigger “false alarms” about an athlete's health, leading to unnecessary changes in training plans. These cascades reduce confidence in AI systems and can compromise decision quality (Sambasivan et al., 2021; Polyzotis et al., 2018).
  3. Bias and Fairness Issues Poor data quality often stems from incomplete datasets that fail to represent diverse athlete populations. When models are trained on limited data—such as metrics from only elite swimmers—they may produce advice that is irrelevant or potentially inappropriate for youth or masters-level athletes. Inclusive and representative data collection is key to mitigating bias (Zhou et al., 2024; Qayyum et al., 2020).
  4. Lack of Standardized Metrics Without standardized methods for recording key metrics (e.g., stroke rate or lap segment times), comparing data across teams or studies becomes difficult. Inconsistent definitions can create confusion when adopting AI solutions, slowing progress and amplifying errors across applications (Priestley et al., 2023).
  5. Data Poisoning and Security Risks When data is poorly managed, it becomes vulnerable to tampering or malicious attacks. In sports, altered performance data could mislead scouts, skew rankings, or even affect betting markets. Implementing robust validation and security measures helps prevent such data poisoning risks (Qayyum et al., 2020).
  6. Resource Constraints and Documentation Issues Under-resourced teams and unclear data collection protocols often lead to avoidable errors. For instance, poorly documented sensor calibration procedures can result in mislabeling data, which later requires extensive effort to correct. Over time, these resource gaps compound inefficiencies (Sambasivan et al., 2021).
  7. Ethical and Legal Challenges Handling sensitive athlete data—including biometric or health-related metrics—requires strict compliance with privacy regulations. Sloppy data management could lead to non-compliance, legal issues, and damage to the trust between athletes and staff (Qayyum et al., 2020; Zhou et al., 2024).
  8. Operational Inefficiencies Poor data quality can significantly slow down progress by requiring constant cleanup and validation. Time spent “firefighting” bad data could be better used to develop advanced training strategies or run additional experiments (Priestley et al., 2023).
  9. Training and Education Gaps Many sports organizations lack proper training in data collection, management, and ethics. Without this foundational knowledge, teams may inadvertently introduce errors into datasets, creating further challenges in scaling AI solutions (Zhou et al., 2024).
  10. Generalization and Representativeness Models trained on narrow datasets often struggle to generalize across different contexts. For example, a model trained exclusively on elite swimmers may offer little value for youth or masters athletes, necessitating additional data collection and retraining (Priestley et al., 2023; Rangineni, 2023).

Poor data quality presents significant challenges for AI adoption in sports. From degraded model performance and ethical risks to operational delays, these pitfalls underscore the need for robust, well-documented, and secure data pipelines. By addressing these challenges, organizations can give coaches, scientists, and support staff a better basis for trusting AI insights—ultimately supporting better training decisions and more equitable outcomes.

Section 3: Core Foundations for Ensuring High-Quality Data in AI/ML

Achieving high-quality data is no accident—it requires intentional strategies and meticulous processes. In sports, especially swimming, data comes from a variety of sources such as lap times, stroke counts, and physiological metrics. To help AI models deliver reliable insights, each data point must be accurate, relevant, and contextually meaningful. Below are the key pillars supporting effective data collection, management, and use.

  1. Intrinsic Data Quality Intrinsic quality focuses on ensuring the data itself is accurate, consistent, and complete. In swimming, even a small inaccuracy—such as a misrecorded lap time—can distort training recommendations and affect athlete outcomes. To achieve high intrinsic quality, sensors like timing pads and wearable devices should undergo regular calibrations. Periodic spot checks, such as comparing automated data with video reviews, help validate the accuracy of key metrics. Automated systems that flag outliers, like stroke rates exceeding physical limits, are also important (Priestley et al., 2023; Rangineni, 2023). These combined measures help keep the data trustworthy enough for AI-assisted analysis.

  2. Contextual Quality Contextual quality ensures that data is relevant, timely, and suitable for its intended AI task. For example, training data gathered from short-course pools may not be applicable to open-water swimming, making segmentation essential. To maintain contextual relevance, teams should clearly define data collection objectives, such as improving starts, turns, or overall endurance. Data should be classified based on conditions like pool size or altitude to provide contextually meaningful insights. Moreover, as training needs evolve, so should data collection processes to keep them aligned with current goals (Priestley et al., 2023; Zhou et al., 2024).

  3. Representational Quality Representational quality focuses on consistent and interpretable data formats across teams and systems. Without standardization, performance data can be misinterpreted—such as when different teams label a 50-meter lap as “50 Free” or “FC_50.” Adopting standardized naming conventions and maintaining a shared data schema across teams help mitigate these issues. Teams should also use metadata to document details about when and how data was collected (Priestley et al., 2023). These measures prevent confusion and improve collaboration between internal and external stakeholders.

  4. Accessibility Accessibility ensures that data is available to authorized users while safeguarding privacy. Coaches, sports scientists, and athletes often need real-time access to performance data to adjust training. Secure cloud-based systems with role-based access control can provide access without compromising security. Additionally, user-friendly dashboards designed for non-technical users allow for broader accessibility. For sensitive athlete data, encryption should be enforced to meet privacy regulations (Zhou et al., 2024; Qayyum et al., 2020). These measures help balance data availability and privacy while supporting effective decision-making.

  5. Data Lifecycle Management Data lifecycle management oversees data from collection to processing, storage, analysis, and eventual archiving or deletion. Traceability is key—without it, errors can be introduced into the AI pipeline unnoticed. Maintaining thorough documentation, including details such as collection dates and sensor calibration logs, helps preserve data integrity. Periodic reviews are essential to remove outdated or irrelevant data while maintaining focus on quality datasets (Rangineni, 2023; Priestley et al., 2023). Backup and disaster recovery strategies further support long-term data reliability.

  6. Ethical and Legal Compliance Ethical and legal compliance is crucial when handling sensitive data, particularly in sports where biometric and health data are involved. Athletes trust that their personal information will be protected and used responsibly. To uphold this trust, teams should anonymize athlete data when possible and ensure that data usage complies with relevant privacy laws and governance requirements. Obtaining informed consent from athletes before collecting and using their data is also essential (Qayyum et al., 2020; Zhou et al., 2024). Failure to adhere to these guidelines risks legal repercussions and reputational harm.

  7. Continuous Monitoring and Improvement Continuous monitoring helps maintain data quality over time as performance data evolves. Swimming programs often introduce new metrics and technologies, making ongoing validation important. Automated validation scripts can detect anomalies, such as unusually short or long lap times, before they affect analyses. Periodic audits help maintain completeness and integrity, while feedback loops involving coaches and athletes allow for the prompt resolution of discrepancies (Bangad et al., 2024; Zhou et al., 2024). This proactive approach helps maintain a dynamic and reliable data pipeline.

  8. Integration of Domain Knowledge Domain knowledge integration leverages the expertise of coaches, sports scientists, and athletes to interpret and validate data effectively. Anomalies, such as a sudden spike in heart rate, may have simple explanations like sensor malfunctions or environmental conditions. Domain experts can distinguish between real issues and equipment errors, preventing unnecessary model adjustments. Collaborating with coaches on data collection protocols and validating AI-driven recommendations against real-world experience improves the chance that insights are useful in practice (Rangineni, 2023; Neutatz et al., 2021). This iterative process helps data-driven decisions stay aligned with practical experience.

By focusing on these core foundations—intrinsic and contextual quality, representational consistency, accessibility, lifecycle management, compliance, continuous monitoring, and domain expertise—organizations can establish more trustworthy data pipelines. For swimming professionals, this supports better-informed training decisions, more accurate athlete feedback, stronger trust, and more useful performance-support workflows.

Summary

In this first part, we have explored the core principles of data quality and shown how poor data can derail even advanced AI projects. Sloppy or incomplete records do more than slow innovation. They can actively mislead coaches, athletes, and analysts. But how do these concepts apply to swimming's current data landscape?

In the next installment, we will examine the practical realities of managing swimming training session data. We will highlight areas where the industry is strong and areas where improvement is needed. We will also discuss the opportunity for a unified framework designed to improve data management across all levels of the sport. Finally, we will answer the key question: Is the swimming fitness and competitive industry data fit for AI? Stay tuned for a closer look at how we can use AI to support better decisions for swimmers at every level.

Note: This article was originally written in English and translated into other languages using automated AI tools so we can share this information with more people. We do our best to keep translations accurate and easy to understand, and we welcome help from the community to improve them. If anything in a translated version is unclear, incorrect, or differs from the English version, the original English text should be considered the official version.

Sources

Bangad, N., Jayaram, V., Sughaturu Krishnappa, M., Banarse, A., Bidkar, D., Nagpal, A., & Parlapalli, V. (2024). A theoretical framework for AI-driven data quality monitoring in high-volume data environments. International Journal of Computer Engineering & Technology, 15, 618-636. https://doi.org/10.5281/zenodo.13878755

Leckey, C., van Dyk, N., Doherty, C., Lawlor, A., & Delahunt, E. (2025). Machine learning approaches to injury risk prediction in sport: A scoping review with evidence synthesis. British Journal of Sports Medicine. https://doi.org/10.1136/bjsports-2024-108576

Neutatz, F., Chen, B., Abedjan, Z., & Wu, E. (2021). From cleaning before ML to cleaning for ML. Bulletin of the IEEE Computer Society Technical Committee on Data Engineering.

Polyzotis, N., Roy, S., Whang, S., & Zinkevich, M. (2018). Data lifecycle challenges in production machine learning: A survey. ACM SIGMOD Record, 47, 17-28. https://doi.org/10.1145/3299887.3299891

Priestley, M., O'Donnell, F., & Simperl, E. (2023). A survey of data quality requirements that matter in ML development pipelines. Journal of Data and Information Quality, 15. https://doi.org/10.1145/3592616

Qayyum, A., Qadir, J., Bilal, M., & Al-Fuqaha, A. (2020). Secure and robust machine learning for healthcare: A survey. IEEE Reviews in Biomedical Engineering. https://doi.org/10.1109/RBME.2020.3013489

Rangineni, S. (2023). An analysis of data quality requirements for machine learning development pipelines frameworks. International Journal of Computer Trends and Technology, 71, 16-27. https://doi.org/10.14445/22312803/IJCTT-V71I8P103

Roh, Y., Heo, G., & Whang, S. E. (2019). A survey on data collection for machine learning: A big data - AI integration perspective. IEEE Transactions on Knowledge and Data Engineering. https://doi.org/10.1109/TKDE.2019.2946162

Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. (2021). "Everyone wants to do the model work, not the data work": Data cascades in high-stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1-15. https://doi.org/10.1145/3411764.3445518

Vec, V., Tomazic, S., Kos, A., & Umek, A. (2024). Trends in real-time artificial intelligence methods in sports: A systematic review. Journal of Big Data. https://doi.org/10.1186/s40537-024-01026-0

Whang, S. E., Roh, Y., Song, H., & Lee, J.-G. (2023). Data collection and quality challenges in deep learning: A data-centric AI perspective. The VLDB Journal, 32. https://doi.org/10.1007/s00778-022-00775-9

Zhou, Y., Tu, F., Sha, K., Ding, J., & Chen, H. (2024). A survey on data quality dimensions and tools for machine learning. 2024 IEEE International Conference on Artificial Intelligence Testing, 120-131. https://doi.org/10.1109/AITest62860.2024.00023

Authors
Diego Torres

Diego Torres

Translators
Wise Racer

Wise Racer


Previous Post
Next Post

Stay up to date with Wise Racer

Subscribe to receive new articles and product updates from Wise Racer. We will send a confirmation email before your subscription is activated.

Email address

​

© 2020 - 2026, Unify Web Solutions Pty Ltd. All rights reserved.