Explore projects

Thèse Guillaume Bernard / Développement / from events to documents / wikivents-projects / wikivents

A Python package to process and represent events from ontologies and semi-structured databases such as Wikidata and Wikipedia.

Archived 0

Updated Oct 30, 2023

Archived 0 0 0

Updated Oct 30, 2023
S

Thèse Guillaume Bernard / Jeux de données / dataset_manipulation_tools / synthesise_ocr_and_segmentation_errors_in_texts

This software enables to damage texts written in any natural language by applying OCR degradation (phantom characters, character degradation, etc.) and by over-segmenting texts (this means splitting regularly the texts in equal parts).

This is useful to reproduce common errors found in historical documents when historical data is missing.

Archived 0

Updated Sep 21, 2022

Archived 0 0 0

Updated Sep 21, 2022
R

Thèse Guillaume Bernard / Développement / from events to documents / request_documents_based_on_events_they_report

Requests to collect documents relating real-world events (themselves described using wikivents) stored in a global index (provided by database_infrastructure_text_mining).

Archived 0

Updated Oct 30, 2023

Archived 0 0 0 0

Updated Oct 30, 2023
N

Thèse Guillaume Bernard / Développement / from documents to events / news_tracking

Command Line Tools to manipulate the document_tracking architecture. It allows to train the Miranda algorithm, to use it and the alternative one, the K-Means implementation. It also provides a tool to evaluate the results.

Archived 0

Updated Sep 21, 2022

Archived 0 0

Updated Sep 21, 2022
D

Thèse Guillaume Bernard / Développement / from documents to events / documents_tracking_resources

Resources and Python API to manipulate datasets of news documents. It manipulates data in the .pickle format with the help of pandas and numpy. It can perform operations on the datasets.

Archived 0

Updated Sep 21, 2022

Archived 0 0 0 0

Updated Sep 21, 2022
D

Thèse Guillaume Bernard / Développement / from documents to events / document_tracking

Implementation of algorithms to detect and track events reported in the news. It provides two alternatives, one supervised, the other unsupervised to track events in the texts.

document ana... document tra... document sim...

Archived 0

Updated Sep 21, 2022

Archived 0 0 0 0

Updated Sep 21, 2022
D

Thèse Guillaume Bernard / Développement / from documents to events / document_processing

Process documents in order to extract tokens, lemmas and named entities from texts. This software depends on spaCy (https://spacy.io/) in order to extract text features and recognise the inner elements.

Archived 0

Updated Aug 31, 2022

Archived 0 0 0 0

Updated Aug 31, 2022
D

Thèse Guillaume Bernard / Développement / from events to documents / database_infrastructure_text_mining

Textual Search Engine Infrastructure based on ElasticSearch (https://www.elastic.co/fr/elasticsearch/) and Lucene (https://lucene.apache.org/). Includes the import scripts to load datasets into the index.

Archived 0

Updated Oct 30, 2023

Archived 0 0 0 0

Updated Oct 30, 2023
C

Thèse Guillaume Bernard / Jeux de données / dataset_manipulation_tools / compute_tf_idf_weights

This software is used to compute TF IDF weighting from texts that are based on the document_tracking_resources format. Vectors and weightings are computed thanks to a resource file that contains a representation of the language used in the same context as the text to weight (news features to weight texts published in the news).

Archived 0

Updated Aug 31, 2022

Archived 0 0 0 0

Updated Aug 31, 2022
C

Thèse Guillaume Bernard / Jeux de données / dataset_manipulation_tools / compute_dense_vectors

This software is used to compute dense vectorisations (sentence embeddings) of sequences of sentences of natural text. It is able to handle multilingual documents until the model used is a multilingual one. This relies on the S-BERT architecture, software and models (https://www.sbert.net/). It computes dense vector representations for tokens, lemmas, entities, etc. of your datasets.

Archived 0

Updated Aug 31, 2022

Archived 0 0 0 0

Updated Aug 31, 2022
A

Thèse Guillaume Bernard / Développement / from events to documents / annotate_events_with_wikidata_identifiers

Annotation tool to check whether the annotation of a Corpus (document_tracking_resources) are correct and truthful.

Archived 0

Updated Aug 30, 2022

Archived 0 0 0 0

Updated Aug 30, 2022