Develop a Data Preprocessing Pipeline for websites, PDF, docx
This issue aims to establish a robust and efficient process for cleaning and preparing the acquired data for use in the chatbot RAG. This involves addressing issues like inconsistencies, missing values, duplicates, and potential errors to ensure high-quality data input. The intention is to use Python libraries to extract and clean the text from source documents.
Edited by Chiara MARGARI