LinkedIn and the use of data for AI: an unclear management of privacy | Festina Lente - Your leading source of AI news | Turtles AI
LinkedIn used user data to train AI models before updating its terms of service, sparking a debate about transparency and user consent. The data collection, while in line with the platform’s policies, raised concerns about privacy practices.
Key Points:
- Data Use for AI Training Not Instantly Transparent
- US users can opt out of the use of their data, but not EU/EEA users.
- LinkedIn only updated its terms of service after the data use was discovered.
- Opt-out consent management raises data protection concerns.
Recent developments have revealed that LinkedIn has begun the process of extracting users’ personal data to train AI models, including those for automatic content generation, without updating its terms of service in a timely manner. This discovery, first reported by 404 Media, has attracted particular attention because the data has been used for purposes new to what was originally stated, raising questions about transparency and user privacy protection. In particular, US users have an option to opt out of the use of their data through a specific function in their profile settings, while European users, protected by strict EU privacy regulations, have not been subject to such processing. Although the opt-out option was already present, LinkedIn’s failure to communicate the policy update before using the data for training has highlighted a flaw in its consent management mechanism. Typically, changes of this magnitude, such as the introduction of the use of personal data to train AI models, are preceded by explicit updates to the terms of use, allowing users to make informed decisions about whether to stay on the platform or change their account settings. However, in this case, LinkedIn only updated its policy after the issue became public, calling into question its adherence to transparency best practices. In the platform’s "Q&A" section, LinkedIn confirmed that it trains its own AI models using data from users, including posts, articles and other interactions, to improve the platform’s functionality. However, it left open the possibility that other models, such as those used to generate content, could be provided by third parties, including its parent company Microsoft.
Specifically, LinkedIn has clarified that the data could be used to optimize services such as post suggestions or personalized recommendations, leveraging user behavioral data. Although LinkedIn has stated that it uses privacy protection techniques, such as removing sensitive information in training data sets, debate remains heated, particularly regarding the appropriateness of opt-out measures. Users can choose to disable the sharing of their data through a specific option in their settings, but this does not cancel any past use of the data. Criticism has been raised by several organizations, including the Open Rights Group (ORG), which has requested the intervention of the UK Information Commissioner’s Office (ICO) to investigate LinkedIn and other platforms that collect user data for training purposes without explicit prior consent. Mariano delli Santi, legal representative of the ORG, highlighted the inadequacy of the opt-out system as a tool to ensure real privacy protection, stressing that the opt-in model should be considered not only a regulatory requirement, but also a common sense measure. The focus has been further heightened by recent moves by other platforms such as Meta, which has restarted its plans to collect data for similar purposes after working with the ICO to simplify the opt-out process. This complex regulatory framework has led Ireland, through its Data Protection Commission (DPC), to closely monitor the situation. LinkedIn has already communicated to the DPC that it is introducing opt-out options for members who do not wish to participate in AI training. However, the DPC itself has stressed that these opt-out settings do not apply to EU and EEA users, as LinkedIn does not currently use their data for such purposes. The use of data to train generative AI models is becoming more widespread among digital platforms, which see significant potential in this data to improve their services. Some, such as Reddit and Tumblr, have started to monetize this data by licensing it to AI developers, creating new forms of economic exploitation of user-generated content. However, this has led to resistance from some users, who in some cases have chosen to delete their content in protest, as happened on Stack Overflow. A more clear-cut closure of the debate is necessary to establish clear boundaries in data management.
The growing demand for data to train AI models continues to raise questions about the implications for user privacy, highlighting the need for greater clarity and transparency from platforms.
