一区二区日本_久久久久久久国产精品_无码国模国产在线观看_久久99深爱久久99精品_亚洲一区二区三区四区五区午夜_日本在线观看一区二区

MLlib is Apache Spark's scalable machine learning library.

Ease of use

Usable in Java, Scala, Python, and R.

MLlib fits into Spark's APIs and interoperates with NumPy in Python (as of Spark 0.9) and R libraries (as of Spark 1.5). You can use any Hadoop data source (e.g. HDFS, HBase, or local files), making it easy to plug into Hadoop workflows.

data = spark.read.format("libsvm")\
  .load("hdfs://...")

model = KMeans(k=10).fit(data)
Calling MLlib in Python

Performance

High-quality algorithms, 100x faster than MapReduce.

Spark excels at iterative computation, enabling MLlib to run fast. At the same time, we care about algorithmic performance: MLlib contains high-quality algorithms that leverage iteration, and can yield better results than the one-pass approximations sometimes used on MapReduce.

Logistic regression in Hadoop and Spark

Runs everywhere

Spark runs on Hadoop, Apache Mesos, Kubernetes, standalone, or in the cloud, against diverse data sources.

You can run Spark using its standalone cluster mode, on EC2, on Hadoop YARN, on Mesos, or on Kubernetes. Access data in HDFS, Apache Cassandra, Apache HBase, Apache Hive, and hundreds of other data sources.

Algorithms

MLlib contains many algorithms and utilities.

ML algorithms include:

  • Classification: logistic regression, naive Bayes,...
  • Regression: generalized linear regression, survival regression,...
  • Decision trees, random forests, and gradient-boosted trees
  • Recommendation: alternating least squares (ALS)
  • Clustering: K-means, Gaussian mixtures (GMMs),...
  • Topic modeling: latent Dirichlet allocation (LDA)
  • Frequent itemsets, association rules, and sequential pattern mining

ML workflow utilities include:

  • Feature transformations: standardization, normalization, hashing,...
  • ML Pipeline construction
  • Model evaluation and hyper-parameter tuning
  • ML persistence: saving and loading models and Pipelines

Other utilities include:

  • Distributed linear algebra: SVD, PCA,...
  • Statistics: summary statistics, hypothesis testing,...

Refer to the MLlib guide for usage examples.

Community

MLlib is developed as part of the Apache Spark project. It thus gets tested and updated with each Spark release.

If you have questions about the library, ask on the Spark mailing lists.

MLlib is still a rapidly growing project and welcomes contributions. If you'd like to submit an algorithm to MLlib, read how to contribute to Spark and send us a patch!

Getting started

To get started with MLlib:

  • Download Spark. MLlib is included as a module.
  • Read the MLlib guide, which includes various usage examples.
  • Learn how to deploy Spark on a cluster if you'd like to run in distributed mode. You can also run locally on a multicore machine without any setup.
主站蜘蛛池模板: 亚洲美女网站 | 欧美在线一区二区三区 | 久久综合99 | 欧美综合久久久 | 一级特黄a大片 | 成人欧美一区二区三区视频xxx | 免费看国产a | 国产精品久久久久久久久久久免费看 | 久久av一区二区三区 | 国产精品综合色区在线观看 | 久久躁日日躁aaaaxxxx | 黄色欧美 | 国产精品精品视频一区二区三区 | 国产一区欧美一区 | 五月天综合网 | 精品久久久999 | 91久久精品一区二区二区 | 91免费视频 | 国产一级片免费看 | 色视频成人在线观看免 | 最新日韩精品 | 国产91在线播放 | 日韩第一页| 少妇久久久久 | av大片| 国产精品成人久久久久 | 亚洲国产精品久久久 | 91免费入口 | 成人在线不卡 | 国产精品久久久久一区二区三区 | 欧美在线a | 欧美美女被c | 日韩三极 | 欧美性生活网 | 国产精品久久久久国产a级 欧美日本韩国一区二区 | 国产一区二区免费 | 成人婷婷 | 久久久性色精品国产免费观看 | 国产精品久久久久一区二区三区 | 国产精品一区在线 | 亚洲嫩草 |