Before using data from public-domain sources to train generative AI models, it is MOST important to confirm that the data-source site:
Content provenance and rights must be established before material is used for training. A site’s demonstrated ownership of its content is the relevant basis for confirming its authority to identify content as public domain or otherwise make it available for use. The U.S. Copyright Office explains that copyright ownership can belong to an author, employer, or transferee, and that public-domain works may be used without permission.
Community Discussion