Python Package for DataHerb: create, search, and load datasets.

Last update: Feb 11, 2022

Overview

The Python Package for DataHerb

A DataHerb Core Service to Create and Load Datasets.

Install

pip install dataherb

Documentation: dataherb.github.io/dataherb-python

The DataHerb Command-Line Tool

Requires Python 3

The DataHerb cli provides tools to create dataset metadata, validate metadata, search dataset in flora, and download dataset.

Search and Download

Search by keyword

dataherb search covid19
# Shows the minimal metadata

Search by dataherb id

dataherb search -i covid19_eu_data
# Shows the full metadata

Download dataset by dataherb id

dataherb download covid19_eu_data
# Downloads this dataset: http://dataherb.io/flora/covid19_eu_data

Create Dataset Using Command Line Tool

We provide a template for dataset creation.

Within a dataset folder where the data files are located, use the following command line tool to create the metadata template.

dataherb create

Upload dataset to remote

Within the dataset folder, run

dataherb upload

UI for all the datasets in a flora

dataherb serve

Use DataHerb in Your Code

Load Data into DataFrame

# Load the package
from dataherb.flora import Flora

# Initialize Flora service
# The Flora service holds all the dataset metadata
use_flora = "path/to/my/flora.json"
dataherb = Flora(flora=use_flora)

# Search datasets with keyword(s)
geo_datasets = dataherb.search("geo")
print(geo_datasets)

# Get a specific file from a dataset and load as DataFrame
tz_df = pd.read_csv(
  dataherb.herb(
      "geonames_timezone"
  ).get_resource(
      "dataset/geonames_timezone.csv"
  )
)
print(tz_df)

The DataHerb Project

What is DataHerb

DataHerb is an open-source data discovery and management tool.

A DataHerb or Herb is a dataset. A dataset comes with the data files, and the metadata of the data files.
A Herb Resource or Resource is a data file in the DataHerb.
A Flora is the combination of all the DataHerbs.

In many data projects, finding the right datasets to enhance your data is one of the most time consuming part. DataHerb adds flavor to your data project. By creating metadata and manage the datasets systematically, locating an dataset is much easier.

Currently, dataherb supports sync dataset between local and S3/git. Each dataset can have its own remote location.

What is DataHerb Flora

We desigined the following workflow to share and index open datasets.

The repo dataherb-flora is a demo flora that lists some datasets and demonstrated on the website https://dataherb.github.io. At this moment, the whole system is being renovated.

Development

Create a conda environment.
Install requirements: pip install -r requirements.txt

Documentation

The source of the documentation for this package is located at docs.

References and Acknolwedgement

dataherb uses datapackage in the core. datapackage is a python library for the data-package standard. The core schema of the dataset is essentially the data-package standard.

Comments

would you like to take a look at our api?

I come across this repo and found it very similar to our API, though much more mature. https://github.com/Glacier-Ice/data-sci-api

we have problems in creating a standard of dataset collection and API documentation for end-users

is there a way we can collaborate?

opened by Stockard 4
Format search results for better ux

The current search result shows too much information. It would be good to format the result into a way that is easier to read and get the id if needed.
enhancement

opened by emptymalei 1
use rapidfuzz instead of fuzzywuzzy

FuzzyWuzzy is GPLv2 licensed which would force you to licence the whole project under GPLv2. I had the same problem on one of my projects and so I wrote rapidfuzz which is implementing the same algorithm but is based on a version of fuzzywuzzy that was MIT Licensed and is therefor MIT Licensed aswell, so it can be used in here without forcing a License change. As a nice bonus it is fully implemented in C++ and comes with a few Algorithmic improvements making it faster than FuzzyWuzzy.

opened by maxbachmann 1
Use One File for Each Herb in Flora
Is it better to have one file for each herb in flora?

Situition

Currently, the flora is defined in a single json file.

It becomes hard to read. This is not fitting into the human-readable principle.

It becomes hard to manage. We are currently sorting everything in the big file. When we have a problem, the whole flora will be unusable.

Solution

Use separate files for herbs.

Simply Copy dataherb.json

Copy dataherb.json to workdir/{id}/dataherb.json or {id}.json will work.

Using folders allows us to put in more files. For example, we can take datapackage content out to make it more managable.

Build the flora from all these files.

[x] Implement this new structure.

Ready for a Demo repo of flora

In this way, we can put up a repo for open datasets easily and allow users to add more easily.

Possible creating process

Create package directly on GitHub by uploading the dataherb.json file.

But there should be a validation process to avoid duplicate id.

[ ] Setup a demo repo as demo flora.

enhancement
opened by emptymalei 0
Overhaul: New Core Management, Local Indexing Webpage, Flexible Flora Database
This is a completely new era of Dataherb.

New Stuff

Supporting S3 as source

Serve whole flora as webpages with search

User config for flora

Multiple flora on one machine

We also redesigned the core.
opened by emptymalei 0
Add dataset using the URL of a remote repo
We don't only upload datasets, we might also want to load datasets from remote.

Here we propose to add the option to add datasets using the URL.

Build a Herb from remote data

Option to add metadata only or download everything.

Adding metadata only will only add data to the flora

Thus we can not find the dataset folder with the corresponding id.

This can be used to decide if a dataset is metadata only or fully downloaded.
opened by emptymalei 0
Sync Flora Metafolder
Managing flora using command line

Version control of the flora is not really hard. We just get into the folder and use git.

But it would be much easier if we can simply run dataherb sync flora

Approaches:

Invoking command line: ref cookiecutter

enhancement
opened by emptymalei 0

Releases(0.1.6)

0.1.6(Feb 10, 2022)
Fixed

Command line tool dataherb configure -l now only opens the folder.

Command line too dataherb download will also display where the dataset is downloaded to. This makes it easier for the user to find the downloaded dataset.

Source code(tar.gz)
Source code(zip)
0.1.5(Aug 12, 2021)

Using Dedicated Folders for Herbs

In the previous versions, we can only use a single file to host all the flora metadata. It will become unmanageable and hard to read as the number of herbs grows. (#14)

In this version, we introduce a new structure for the flora metadata. Each herb is getting its own folder! This structure makes it easier for us to read and manage by hand. It is also better for version-controling your flora.

(🌱 Best wishes to your herbs in their own pots. )
Source code(tar.gz)
Source code(zip)
0.1.4(Aug 7, 2021)
Added

🎉 Better search result formatting in terminal (See docs for a screenshot.)

📺 Show config using dataherb configure --show

Changed

Better config management. Config has been promoted to a class.

Source code(tar.gz)
Source code(zip)
0.1.3(Aug 7, 2021)
Added

Server to serve flora as a website

Configuration system

Remove herb from flora

Add herb to flora

and more

Source code(tar.gz)
Source code(zip)
0.0.5(Mar 14, 2020)
Now we can use

dataherb validate

to validate the metadata file.
Source code(tar.gz)
Source code(zip)
0.0.3(Feb 23, 2020)

dataherb command line tool now automatically finds the data files and generate part of the metadata based on the files. CSV files are automatically parsed.
Source code(tar.gz)
Source code(zip)
0.0.2(Feb 16, 2020)

Source code(tar.gz)
Source code(zip)

Owner

DataHerb

Get datasets in a blink of an eye | Experimenting with simple modular small dataset discovery

GitHub Repository https://dataherb.github.io/dataherb-python

Codes for the collection and predictive processing of bitcoin from the API of coinmarketcap

5 Apr 26, 2022

Python Library for learning (Structure and Parameter) and inference (Statistical and Causal) in Bayesian Networks.

pgmpy pgmpy is a python library for working with Probabilistic Graphical Models. Documentation and list of algorithms supported is at our official sit

2.2k Dec 25, 2022

Extract data from a wide range of Internet sources into a pandas DataFrame.

pandas-datareader Up to date remote data access for pandas, works for multiple versions of pandas. Installation Install using pip pip install pandas-d

2.5k Jan 09, 2023

Udacity-api-reporting-pipeline - Udacity api reporting pipeline

udacity-api-reporting-pipeline In this exercise, you'll use portions of each of

1 Feb 15, 2022

A utility for functional piping in Python that allows you to access any function in any scope as a partial.

WithPartial Introduction WithPartial is a simple utility for functional piping in Python. The package exposes a context manager (used with with) calle

1 Oct 26, 2021

Open source platform for Data Science Management automation

Hydrosphere examples This repo contains demo scenarios and pre-trained models to show Hydrosphere capabilities. Data and artifacts management Some mod

6 Aug 10, 2021

Using Python to derive insights on particular Pokemon, Types, Generations, and Stats

Pokémon Analysis Andreas Nikolaidis February 2022 Introduction Exploratory Analysis Correlations & Descriptive Statistics Principal Component Analysis

1 Feb 18, 2022

Statsmodels: statistical modeling and econometrics in Python

About statsmodels statsmodels is a Python package that provides a complement to scipy for statistical computations including descriptive statistics an

8k Dec 29, 2022

International Space Station data with Python research 🌎

International Space Station data with Python research 🌎 Plotting ISS trajectory, calculating the velocity over the earth and more. Plotting trajector

41 Jun 16, 2022

CubingB is a timer/analyzer for speedsolving Rubik's cubes, with smart cube support

CubingB is a timer/analyzer for speedsolving Rubik's cubes (and related puzzles). It focuses on supporting "smart cubes" (i.e. bluetooth cubes) for recording the exact moves of a solve in real time.

5 Sep 18, 2022

Minimal working example of data acquisition with nidaqmx python API

Data Aquisition using NI-DAQmx python API Based on this project It is a minimal working example for data acquisition using the NI-DAQmx python API. It

1 Nov 05, 2021

Fast, flexible and easy to use probabilistic modelling in Python.

Please consider citing the JMLR-MLOSS Manuscript if you've used pomegranate in your academic work! pomegranate is a package for building probabilistic

3k Jan 02, 2023

Stock Analysis dashboard Using Streamlit and Python

StDashApp Stock Analysis Dashboard Using Streamlit and Python If you found the content useful and want to support my work, you can buy me a coffee! Th

27 Dec 09, 2022

PostQF is a user-friendly Postfix queue data filter which operates on data produced by postqueue -j.

11 Nov 24, 2022

MotorcycleParts DataAnalysis python

We work with the accounting department of a company that sells motorcycle parts. The company operates three warehouses in a large metropolitan area.

1 Jan 12, 2022

Useful tool for inserting DataFrames into the Excel sheet.

PyCellFrame Insert Pandas DataFrames into the Excel sheet with a bunch of conditions Install pip install pycellframe Usage Examples Let's suppose that

1 Feb 16, 2022

Desafio 1 ~ Bantotal

Challenge 01 | Bantotal Please read the instructions for the challenge by selecting your preferred language below: Español Português License Copyright

44 Sep 28, 2022

Processo de ETL (extração, transformação, carregamento) realizado pela equipe no projeto final do curso da Soul Code Academy.

1 Feb 03, 2022

4CAT: Capture and Analysis Toolkit

4CAT: Capture and Analysis Toolkit 4CAT is a research tool that can be used to analyse and process data from online social platforms. Its goal is to m

147 Dec 20, 2022

Data-sets from the survey and analysis

bachelor-thesis "Umfragewerte.xlsx" contains the orginal survey results. "umfrage_alle.csv" contains the survey results but one participant is cancele

1 Jan 26, 2022

Python Package for DataHerb: create, search, and load datasets.

Related tags

Overview

The Python Package for DataHerb

A DataHerb Core Service to Create and Load Datasets.

Install

The DataHerb Command-Line Tool

Search and Download

Create Dataset Using Command Line Tool

Upload dataset to remote

UI for all the datasets in a flora

Use DataHerb in Your Code

Load Data into DataFrame

The DataHerb Project

What is DataHerb

What is DataHerb Flora

Development

Documentation

References and Acknolwedgement

Comments

would you like to take a look at our api?

Format search results for better ux

use rapidfuzz instead of fuzzywuzzy

Use One File for Each Herb in Flora

Situition

Solution

Simply Copy dataherb.json

Ready for a Demo repo of flora

Overhaul: New Core Management, Local Indexing Webpage, Flexible Flora Database

New Stuff

Add dataset using the URL of a remote repo

Sync Flora Metafolder

Managing flora using command line

Releases(0.1.6)

0.1.6(Feb 10, 2022)

Fixed

0.1.5(Aug 12, 2021)

Using Dedicated Folders for Herbs

0.1.4(Aug 7, 2021)

Added

Changed

0.1.3(Aug 7, 2021)