Popularity

7.5

Stable

Activity

0.0

Stable

Stars 1,112

Watchers 39

Forks 93

Last Commit about 2 years ago

Programming language: Go

License: GNU General Public License v3.0 or later

Tags: Science And Data Analysis

dataframe-go alternatives and similar packages

Based on the "Science and Data Analysis" category.
Alternatively, view dataframe-go alternatives based on common mentions on social networks and blogs.

gonum

9.3 8.2 dataframe-go VS gonum

Gonum is a set of numeric libraries for the Go programming language. It contains libraries for matrices, statistics, optimization, and more
Stats

8.5 2.0 dataframe-go VS Stats

A well tested and comprehensive Golang statistics library package with no dependencies.

WorkOS - The modern identity platform for B2B SaaS

The APIs are flexible and easy-to-use, supporting authentication, user identity, and complex enterprise features like SSO and SCIM provisioning.

Promo workos.com

gonum/plot

8.4 5.4 dataframe-go VS gonum/plot

A repository for plotting and visualizing data
gosl

8.1 6.0 dataframe-go VS gosl

Linear algebra, eigenvalues, FFT, Bessel, elliptic, orthogonal polys, geometry, NURBS, numerical quadrature, 3D transfinite interpolation, random numbers, Mersenne twister, probability distributions, optimisation, differential equations.
streamtools

7.7 0.0 dataframe-go VS streamtools

tools for working with streams of data
chart

7.2 0.0 dataframe-go VS chart

Provide basic charts in go
goraph

7.2 0.0 dataframe-go VS goraph

Package goraph implements graph data structure and algorithms.
go-dsp

7.2 0.0 dataframe-go VS go-dsp

Digital Signal Processing for Go
graph

6.9 0.0 dataframe-go VS graph

Graph algorithms and data structures
gonum/mat64

6.8 0.0 dataframe-go VS gonum/mat64

DISCONTINUED. The general purpose package for matrix computation. Package mat64 provides basic linear algebra operations for float64 matrices.
ewma

6.3 0.0 dataframe-go VS ewma

Exponentially Weighted Moving Average algorithms for Go.
go.matrix

6.2 0.0 dataframe-go VS go.matrix

linear algebra for go
calendarheatmap

6.0 6.4 dataframe-go VS calendarheatmap

📅 Calendar heatmap inspired by GitHub contribution activity
gohistogram

5.4 0.0 dataframe-go VS gohistogram

DISCONTINUED. Streaming approximate histograms in Go
TextRank

5.2 0.0 dataframe-go VS TextRank

:wink: :cyclone: :strawberry: TextRank implementation in Golang with extendable features (summarization, phrase extraction) and multithreading (goroutine).
blas

4.9 0.0 dataframe-go VS blas

Go implementation of BLAS (Basic Linear Algebra Subprograms)
sparse

4.9 0.0 dataframe-go VS sparse

Sparse matrix formats for linear algebra supporting scientific and machine learning applications
pagerank

4.3 0.0 dataframe-go VS pagerank

Weighted PageRank implementation in Go
go-estimate

4.2 3.2 dataframe-go VS go-estimate

State estimation and filtering algorithms in Go
geom

3.7 0.0 dataframe-go VS geom

2d geometry for golang
vectormath

3.5 0.0 dataframe-go VS vectormath

Vectormath for Go
evaler

3.5 0.0 dataframe-go VS evaler

Implements a simple floating point arithmetic expression evaluator in Go (golang).
jsonl-graph

3.5 0.0 dataframe-go VS jsonl-graph

🏝 JSONL Graph Tools
gostat

2.9 0.0 dataframe-go VS gostat

Collection of statistical routines in golang
gograph

2.8 7.2 dataframe-go VS gograph

A golang generic graph library that provides mathematical graph-theory and algorithms.
triangolatte

2.6 0.0 dataframe-go VS triangolatte

2D triangulation library. Allows translating lines and polygons (both based on points) to the language of GPUs.
permutation

2.5 0.0 dataframe-go VS permutation

Simple permutation package for golang
piecewiselinear

2.3 4.5 dataframe-go VS piecewiselinear

tiny linear interpolation library for go
goent

2.3 0.0 dataframe-go VS goent

GO Implementation of Entropy Measures
ode

2.1 0.0 dataframe-go VS ode

An ordinary differential equation solving library in golang.
PiHex

2.0 0.0 dataframe-go VS PiHex

PiHex Library, written in Go, generates a hexadecimal number sequence in the number Pi in the range from 0 to 10,000,000.
GoStats

1.8 0.0 dataframe-go VS GoStats

GoStats is a go library for math statistics mostly used in ML domains, it covers most of the statistical measures functions.
rootfinding

1.6 0.0 dataframe-go VS rootfinding

root-finding library
godesim

1.6 0.0 dataframe-go VS godesim

ODE system solver made simple. For IVPs (initial value problems).
assocentity

1.5 4.1 dataframe-go VS assocentity

Package assocentity returns the mean distance from tokens to an entity and its synonyms
gofrac

1.3 0.0 dataframe-go VS gofrac

A fractions library for go (http://golang.org)
bradleyterry

1.2 0.0 dataframe-go VS bradleyterry

Package to do Bradley-Terry Model pairwise compairsons
go-fn

1.2 0.0 dataframe-go VS go-fn

Automatically exported from code.google.com/p/go-fn
go-gt

0.8 0.0 dataframe-go VS go-gt

Automatically exported from code.google.com/p/go-gt
gobbs

0.6 0.0 dataframe-go VS gobbs

DISCONTINUED. A Blum-Blum-Shub-Generator written in Go
gocomplex

0.3 0.0 dataframe-go VS gocomplex

Automatically exported from code.google.com/p/gocomplex
mudlark-go

0.1 0.0 dataframe-go VS mudlark-go

DISCONTINUED. A collection of packages providing (hopefully) useful code for use in software using Google's Go programming language.

Do you think we are missing an alternative of dataframe-go or a related project?

Add another 'Science and Data Analysis' Package

Popular Comparisons

README

⭐ the project to show your appreciation. :arrow_upper_right:

Dataframes are used for statistics, machine-learning, and data manipulation/exploration. You can think of a Dataframe as an excel spreadsheet. This package is designed to be light-weight and intuitive.

⚠️ The package is production ready but the API is not stable yet. Once Go 1.18 (Generics) is introduced, the ENTIRE package will be rewritten. For example, there will only be 1 generic Series type. After that, version 1.0.0 will be tagged.

It is recommended your package manager locks to a commit id instead of the master branch directly. ⚠️

Features

Importing from CSV, JSONL, Parquet, MySQL & PostgreSQL
Exporting to CSV, JSONL, Excel, Parquet, MySQL & PostgreSQL
Developer Friendly
Flexible - Create custom Series (custom data types)
Performant
Interoperability with gonum package.
pandas sub-package
Fake data generation
Interpolation (ForwardFill, BackwardFill, Linear, Spline, Lagrange)
Time-series Forecasting (SES, Holt-Winters)
Math functions
Plotting (cross-platform)

See Tutorial here.

Installation

go get -u github.com/rocketlaunchr/dataframe-go

import dataframe "github.com/rocketlaunchr/dataframe-go"

DataFrames

Creating a DataFrame


s1 := dataframe.NewSeriesInt64("day", nil, 1, 2, 3, 4, 5, 6, 7, 8)
s2 := dataframe.NewSeriesFloat64("sales", nil, 50.3, 23.4, 56.2, nil, nil, 84.2, 72, 89)
df := dataframe.NewDataFrame(s1, s2)

fmt.Print(df.Table())

OUTPUT:
+-----+-------+---------+
|     |  DAY  |  SALES  |
+-----+-------+---------+
| 0:  |   1   |  50.3   |
| 1:  |   2   |  23.4   |
| 2:  |   3   |  56.2   |
| 3:  |   4   |   NaN   |
| 4:  |   5   |   NaN   |
| 5:  |   6   |  84.2   |
| 6:  |   7   |   72    |
| 7:  |   8   |   89    |
+-----+-------+---------+
| 8X2 | INT64 | FLOAT64 |
+-----+-------+---------+

Insert and Remove Row


df.Append(nil, 9, 123.6)

df.Append(nil, map[string]interface{}{
    "day":   10,
    "sales": nil,
})

df.Remove(0)

OUTPUT:
+-----+-------+---------+
|     |  DAY  |  SALES  |
+-----+-------+---------+
| 0:  |   2   |  23.4   |
| 1:  |   3   |  56.2   |
| 2:  |   4   |   NaN   |
| 3:  |   5   |   NaN   |
| 4:  |   6   |  84.2   |
| 5:  |   7   |   72    |
| 6:  |   8   |   89    |
| 7:  |   9   |  123.6  |
| 8:  |  10   |   NaN   |
+-----+-------+---------+
| 9X2 | INT64 | FLOAT64 |
+-----+-------+---------+

Update Row


df.UpdateRow(0, nil, map[string]interface{}{
    "day":   3,
    "sales": 45,
})

Sorting


sks := []dataframe.SortKey{
    {Key: "sales", Desc: true},
    {Key: "day", Desc: true},
}

df.Sort(ctx, sks)

OUTPUT:
+-----+-------+---------+
|     |  DAY  |  SALES  |
+-----+-------+---------+
| 0:  |   9   |  123.6  |
| 1:  |   8   |   89    |
| 2:  |   6   |  84.2   |
| 3:  |   7   |   72    |
| 4:  |   3   |  56.2   |
| 5:  |   2   |  23.4   |
| 6:  |  10   |   NaN   |
| 7:  |   5   |   NaN   |
| 8:  |   4   |   NaN   |
+-----+-------+---------+
| 9X2 | INT64 | FLOAT64 |
+-----+-------+---------+

Iterating

You can change the step and starting row. It may be wise to lock the DataFrame before iterating.

The returned value is a map containing the name of the series (string) and the index of the series (int) as keys.


iterator := df.ValuesIterator(dataframe.ValuesOptions{0, 1, true}) // Don't apply read lock because we are write locking from outside.

df.Lock()
for {
    row, vals, _ := iterator()
    if row == nil {
        break
    }
    fmt.Println(*row, vals)
}
df.Unlock()

OUTPUT:
0 map[day:1 0:1 sales:50.3 1:50.3]
1 map[sales:23.4 1:23.4 day:2 0:2]
2 map[day:3 0:3 sales:56.2 1:56.2]
3 map[1:<nil> day:4 0:4 sales:<nil>]
4 map[day:5 0:5 sales:<nil> 1:<nil>]
5 map[sales:84.2 1:84.2 day:6 0:6]
6 map[day:7 0:7 sales:72 1:72]
7 map[day:8 0:8 sales:89 1:89]

Statistics

You can easily calculate statistics for a Series using the gonum or montanaflynn/stats package.

SeriesFloat64 and SeriesTime provide access to the exported Values field to seamlessly interoperate with external math-based packages.

Example

Some series provide easy conversion using the ToSeriesFloat64 method.

import "gonum.org/v1/gonum/stat"

s := dataframe.NewSeriesInt64("random", nil, 1, 2, 3, 4, 5, 6, 7, 8)
sf, _ := s.ToSeriesFloat64(ctx)

Mean

mean := stat.Mean(sf.Values, nil)

Median

import "github.com/montanaflynn/stats"
median, _ := stats.Median(sf.Values)

Standard Deviation

std := stat.StdDev(sf.Values, nil)

Plotting (cross-platform)

import (
    chart "github.com/wcharczuk/go-chart"
    "github.com/rocketlaunchr/dataframe-go/plot"
    wc "github.com/rocketlaunchr/dataframe-go/plot/wcharczuk/go-chart"
)

sales := dataframe.NewSeriesFloat64("sales", nil, 50.3, nil, 23.4, 56.2, 89, 32, 84.2, 72, 89)
cs, _ := wc.S(ctx, sales, nil, nil)

graph := chart.Chart{Series: []chart.Series{cs}}

plt, _ := plot.Open("Monthly sales", 450, 300)
graph.Render(chart.SVG, plt)
plt.Display(plot.None)
<-plt.Closed

Output:

Math Functions

import "github.com/rocketlaunchr/dataframe-go/math/funcs"

res := 24
sx := dataframe.NewSeriesFloat64("x", nil, utils.Float64Seq(1, float64(res), 1))
sy := dataframe.NewSeriesFloat64("y", &dataframe.SeriesInit{Size: res})
df := dataframe.NewDataFrame(sx, sy)

fn := funcs.RegFunc("sin(2*𝜋*x/24)")
funcs.Evaluate(ctx, df, fn, 1)

Output:

Importing Data

The imports sub-package has support for importing csv, jsonl, parquet, and directly from a SQL database. The DictateDataType option can be set to specify the true underlying data type. Alternatively, InferDataTypes option can be set.

CSV

csvStr := `
Country,Date,Age,Amount,Id
"United States",2012-02-01,50,112.1,01234
"United States",2012-02-01,32,321.31,54320
"United Kingdom",2012-02-01,17,18.2,12345
"United States",2012-02-01,32,321.31,54320
"United Kingdom",2012-05-07,NA,18.2,12345
"United States",2012-02-01,32,321.31,54320
"United States",2012-02-01,32,321.31,54320
Spain,2012-02-01,66,555.42,00241
`
df, err := imports.LoadFromCSV(ctx, strings.NewReader(csvStr))

OUTPUT:
+-----+----------------+------------+-------+---------+-------+
|     |    COUNTRY     |    DATE    |  AGE  | AMOUNT  |  ID   |
+-----+----------------+------------+-------+---------+-------+
| 0:  | United States  | 2012-02-01 |  50   |  112.1  | 1234  |
| 1:  | United States  | 2012-02-01 |  32   | 321.31  | 54320 |
| 2:  | United Kingdom | 2012-02-01 |  17   |  18.2   | 12345 |
| 3:  | United States  | 2012-02-01 |  32   | 321.31  | 54320 |
| 4:  | United Kingdom | 2015-05-07 |  NaN  |  18.2   | 12345 |
| 5:  | United States  | 2012-02-01 |  32   | 321.31  | 54320 |
| 6:  | United States  | 2012-02-01 |  32   | 321.31  | 54320 |
| 7:  |     Spain      | 2012-02-01 |  66   | 555.42  |  241  |
+-----+----------------+------------+-------+---------+-------+
| 8X5 |     STRING     |    TIME    | INT64 | FLOAT64 | INT64 |
+-----+----------------+------------+-------+---------+-------+

Exporting Data

The exports sub-package has support for exporting to csv, jsonl, parquet, Excel and directly to a SQL database.

Optimizations

If you know the number of rows in advance, you can set the capacity of the underlying slice of a series using SeriesInit{}. This will preallocate memory and provide speed improvements.

Generic Series

Out of the box, there is support for string, time.Time, float64 and int64. Automatic support exists for float32 and all types of integers. There is a convenience function provided for dealing with bool. There is also support for complex128 inside the xseries subpackage.

There may be times that you want to use your own custom data types. You can either implement your own Series type (more performant) or use the Generic Series (more convenient).

civil.Date

import "time"
import "cloud.google.com/go/civil"

sg := dataframe.NewSeriesGeneric("date", civil.Date{}, nil, civil.Date{2018, time.May, 01}, civil.Date{2018, time.May, 02}, civil.Date{2018, time.May, 03})
s2 := dataframe.NewSeriesFloat64("sales", nil, 50.3, 23.4, 56.2)

df := dataframe.NewDataFrame(sg, s2)

OUTPUT:
+-----+------------+---------+
|     |    DATE    |  SALES  |
+-----+------------+---------+
| 0:  | 2018-05-01 |  50.3   |
| 1:  | 2018-05-02 |  23.4   |
| 2:  | 2018-05-03 |  56.2   |
+-----+------------+---------+
| 3X2 | CIVIL DATE | FLOAT64 |
+-----+------------+---------+

Tutorial

Create some fake data

Let's create a list of 8 "fake" employees with a name, title and base hourly wage rate.

import "golang.org/x/exp/rand"
import "rocketlaunchr/dataframe-go/utils/faker"

src := rand.NewSource(uint64(time.Now().UTC().UnixNano()))
df := faker.NewDataFrame(8, src, faker.S("name", 0, "Name"), faker.S("title", 0.5, "JobTitle"), faker.S("base rate", 0, "Number", 15, 50))

+-----+----------------+----------------+-----------+
|     |      NAME      |     TITLE      | BASE RATE |
+-----+----------------+----------------+-----------+
| 0:  | Cordia Jacobi  |   Consultant   |    42     |
| 1:  | Nickolas Emard |      NaN       |    22     |
| 2:  | Hollis Dickens | Representative |    22     |
| 3:  | Stacy Dietrich |      NaN       |    43     |
| 4:  |  Aleen Legros  |    Officer     |    21     |
| 5:  |  Adelia Metz   |   Architect    |    18     |
| 6:  | Sunny Gerlach  |      NaN       |    28     |
| 7:  | Austin Hackett |      NaN       |    39     |
+-----+----------------+----------------+-----------+
| 8X3 |     STRING     |     STRING     |   INT64   |
+-----+----------------+----------------+-----------+

Apply Function

Let's give a promotion to everyone by doubling their salary.

s := df.Series[2]

applyFn := dataframe.ApplySeriesFn(func(val interface{}, row, nRows int) interface{} {
    return 2 * val.(int64)
})

dataframe.Apply(ctx, s, applyFn, dataframe.FilterOptions{InPlace: true})

+-----+----------------+----------------+-----------+
|     |      NAME      |     TITLE      | BASE RATE |
+-----+----------------+----------------+-----------+
| 0:  | Cordia Jacobi  |   Consultant   |    84     |
| 1:  | Nickolas Emard |      NaN       |    44     |
| 2:  | Hollis Dickens | Representative |    44     |
| 3:  | Stacy Dietrich |      NaN       |    86     |
| 4:  |  Aleen Legros  |    Officer     |    42     |
| 5:  |  Adelia Metz   |   Architect    |    36     |
| 6:  | Sunny Gerlach  |      NaN       |    56     |
| 7:  | Austin Hackett |      NaN       |    78     |
+-----+----------------+----------------+-----------+
| 8X3 |     STRING     |     STRING     |   INT64   |
+-----+----------------+----------------+-----------+

Create a Time series

Let's inform all employees separately on sequential days.

import "rocketlaunchr/dataframe-go/utils/utime"

mts, _ := utime.NewSeriesTime(ctx, "meeting time", "1D", time.Now().UTC(), false, utime.NewSeriesTimeOptions{Size: &[]int{8}[0]})
df.AddSeries(mts, nil)

+-----+----------------+----------------+-----------+--------------------------------+
|     |      NAME      |     TITLE      | BASE RATE |          MEETING TIME          |
+-----+----------------+----------------+-----------+--------------------------------+
| 0:  | Cordia Jacobi  |   Consultant   |    84     |   2020-02-02 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 1:  | Nickolas Emard |      NaN       |    44     |   2020-02-03 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 2:  | Hollis Dickens | Representative |    44     |   2020-02-04 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 3:  | Stacy Dietrich |      NaN       |    86     |   2020-02-05 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 4:  |  Aleen Legros  |    Officer     |    42     |   2020-02-06 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 5:  |  Adelia Metz   |   Architect    |    36     |   2020-02-07 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 6:  | Sunny Gerlach  |      NaN       |    56     |   2020-02-08 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 7:  | Austin Hackett |      NaN       |    78     |   2020-02-09 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
+-----+----------------+----------------+-----------+--------------------------------+
| 8X4 |     STRING     |     STRING     |   INT64   |              TIME              |
+-----+----------------+----------------+-----------+--------------------------------+

Filtering

Let's filter out our senior employees (they have titles) for no reason.

filterFn := dataframe.FilterDataFrameFn(func(vals map[interface{}]interface{}, row, nRows int) (dataframe.FilterAction, error) {
    if vals["title"] == nil {
        return dataframe.DROP, nil
    }
    return dataframe.KEEP, nil
})

seniors, _ := dataframe.Filter(ctx, df, filterFn)

+-----+----------------+----------------+-----------+--------------------------------+
|     |      NAME      |     TITLE      | BASE RATE |          MEETING TIME          |
+-----+----------------+----------------+-----------+--------------------------------+
| 0:  | Cordia Jacobi  |   Consultant   |    84     |   2020-02-02 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 1:  | Hollis Dickens | Representative |    44     |   2020-02-04 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 2:  |  Aleen Legros  |    Officer     |    42     |   2020-02-06 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
| 3:  |  Adelia Metz   |   Architect    |    36     |   2020-02-07 23:13:53.015324   |
|     |                |                |           |           +0000 UTC            |
+-----+----------------+----------------+-----------+--------------------------------+
| 4X4 |     STRING     |     STRING     |   INT64   |              TIME              |
+-----+----------------+----------------+-----------+--------------------------------+

Other useful packages

awesome-svelte - Resources for killing react
dbq - Zero boilerplate database operations for Go
electron-alert - SweetAlert2 for Electron Applications
google-search - Scrape google search results
igo - A Go transpiler with cool new syntax such as fordefer (defer for for-loops)
mysql-go - Properly cancel slow MySQL queries
react - Build front end applications using Go
remember-go - Cache slow database queries
testing-go - Testing framework for unit testing

Legal Information

The license is a modified MIT license. Refer to LICENSE file for more details.

*Note that all licence references and agreements mentioned in the dataframe-go README section above are relevant to that project's source code only.

dataframe-go

DataFrames for Go: For statistics, machine-learning, and data manipulation/exploration