Added samples: GitHubLabeler and GettingStarted #198

OliaG · 2018-05-22T05:51:57Z

Iris sample
Sentiment Analysis sample
TaxiFare sample
GitHub issues classification sample

- Iris sample - Sentiment Analysis sample - TaxiFare sample -GitHub issues classification sample

asthana86

👍

Would recommend Adding a few comments around training and evaluation pipelines for the iris, taxi and github samples as well similar to the sentiment analysis one.

justinormont · 2018-05-22T07:13:38Z

Samples/GettingStarted/BinaryClassification_SentimentAnalysis/Program.cs

+
+            //add a FastTreeBinaryClassifier, the decision tree learner for this project, and 
+            //three hyperparameters to be used for tuning decision tree performance 
+            pipeline.Add(new FastTreeBinaryClassifier() {NumLeaves = 5, NumTrees = 5, MinDocumentsInLeafs = 2});


I'd recommend Bigrams+Trichargrams w/ AveragedPerceptronBinaryClassifier{iter=10} for text. AveragedPerceptron generally wins on text vs. FastTree in terms of accuracy & speed.

Thank you! I will use them for advance scenario samples. This one is a "Hello world" type of sample where I'm trying to make the pipeline as simple as possible. And in the next samples I'll demonstrate how the results can be improved with different transforms.

justinormont · 2018-05-22T07:16:59Z

Samples/GettingStarted/BinaryClassification_SentimentAnalysis/Program.cs

+
+            // TextFeaturizer is a transform that will be used to featurize an input column. 
+            // This is used to format and clean the data.
+            pipeline.Add(new TextFeaturizer("Features", "SentimentText"));


Would be a good opportunity to show off some of the hyperparameters of the TextFeaturizer.

pipeline.Add(new TextFeaturizer("Features", "SentimentText") { WordFeatureExtractor = new NGramNgramExtractor() { NgramLength = 2, AllLengths = true }, CharFeatureExtractor = new NGramNgramExtractor() { NgramLength = 3, AllLengths = false } });

Others of course are available:

machinelearning/test/Microsoft.ML.Tests/Scenarios/SentimentPredictionTests.cs

Line 54 in 86f5ee6

pipeline.Add(new TextFeaturizer("Features", "SentimentText")

This makes sense in other samples we are adding or going to add. For getting started sample the easier the better.

justinormont · 2018-05-22T07:22:38Z

Samples/GettingStarted/MulticlassClassification_Iris/Program.cs

+            Console.WriteLine($"    AccuracyMacro = {metrics.AccuracyMacro:0.####}, a value between 0 and 1, the closer to 1, the better");
+            Console.WriteLine($"    AccuracyMicro = {metrics.AccuracyMicro:0.####}, a value between 0 and 1, the closer to 1, the better");
+            Console.WriteLine($"    LogLoss = {metrics.LogLoss:0.####}, the closer to 0, the better");
+            Console.WriteLine($"    LogLoss for class 1 = {metrics.PerClassLogLoss[0]:0.####}, the closer to 0, the better");


Would it be possible to get the actual name for class 1? When a user runs on their own dataset, they will likely be interested in their actual class names.

shauheen · 2018-05-22T14:06:41Z

This PR closes #140 @OliaG

markusweimer

I did not review the samples in detail. However, I suggest some overall changes:

It would be good to have XML Docs on the major classes and methods in each project.
Where will the long form description of each sample go? My vote would be for a README.md in each of the project folders.

markusweimer · 2018-05-22T14:10:18Z

Samples/GitHubLabeler/GitHubLabeler/App.config

+<?xml version="1.0" encoding="utf-8" ?>
+<configuration>
+  <appSettings>
+    <!--TODO: Please enter your own credentials here. -->


We should add a warning to this such that people know that they should not commit this file to a public repo, once changed.

How far would adding this to the .gitignore go for helping folks to not commit this back to a repo? I've done regex checking in githooks before (eg: /(password|access_token) ?=/i), but I don't think githooks can be added to a repo & auto-run by other users.

I don't think there is a great solution for this. In general, the recommendation is to avoid using credentials from code and instead use certificates, user-based authentication, or a service like Azure KeyVault. Unfortunately this doesn't work for samples very well...

TomFinley · 2018-05-22T14:47:44Z

Samples/GitHubLabeler/.gitignore

@@ -0,0 +1,288 @@
+## Ignore Visual Studio temporary files, build results, and
+## files generated by popular Visual Studio add-ons.
+##


I don't think checking in this .gitignore file was intentional, was it?

Agreed. We should only use the one from the root.

TomFinley · 2018-05-22T14:50:58Z

Samples/GettingStarted/Regression_TaxiFarePrediction/Program.cs

+    {
+        private static string AppPath => Path.GetDirectoryName(Environment.GetCommandLineArgs()[0]);
+        private static string TrainDataPath => Path.Combine(AppPath, @"..\..\..\..\datasets\", "taxi-fare-train.csv");
+        private static string TestDataPath => Path.Combine(AppPath, @"..\..\..\..\datasets\", "taxi-fare-test.csv");


We'd like these samples to work in multiple places, I'd think, including cross platform. Do these \ style paths work in Linux?

No, they don't work on Linux.

Instead, we should use Path.Combine("..", "..", "..", "..", "datasets").

TomFinley · 2018-05-22T14:53:11Z

Samples/GettingStarted/BinaryClassification_SentimentAnalysis/Program.cs

+        private static string AppPath => Path.GetDirectoryName(Environment.GetCommandLineArgs()[0]);
+        private static string TrainDataPath => Path.Combine(AppPath, @"..\..\..\..\datasets\", "imdb_labelled.txt");
+        private static string TestDataPath => Path.Combine(AppPath, @"..\..\..\..\datasets\", "yelp_labelled.txt");
+        private static string ModelPath => Path.Combine(AppPath, "Models", "SentimentModel.zip");


"Models", "SentimentModel.zip"); [](start = 65, length = 32)

I've noticed you have I think unintentionally also put the model files (the .zips) in your PR. I think that was unintentional, and is just an artifafct of the run. Either way we have the mechanism to create this artifact here, and we generally try to avoid checking in binary artifacts if it can be helped. Could you remove them?

TomFinley · 2018-05-22T14:58:14Z

Samples/GitHubLabeler/GitHubLabeler/Program.cs

+    {
+        private static async Task Main(string[] args)
+        {
+            if (args.Length != 1)


Is it possible to simplify this sample a bit, so that people don't have to learn some syntax of a simple command line program? As far as I see the sample should just be a straight train of a model, followed by a label, right? Or are we showing how the artifact is preserved across trainings?

Yes, the intention was to demonstrate how it is used in real life where you train it just once and inference many times. Completely agree that for the tutorial samples we want them to be simple train-evaluate-predict. This one is an end-to-end app that you would use for labeling issues. I will probably move it to separate folder for "apps infused with ML" examples.

TomFinley · 2018-05-22T14:59:30Z

Samples/GitHubLabeler/README.md

+## Labeling
+When the model is trained, it can be used for predicting new issue's label. To do so, run the application with `"label"` key:
+```
+C:\GitHubLabeler\GitHubLabeler\bin\Debug\netcoreapp2.0>dotnet GitHubLabeler.dll label


C:\GitHubLabeler\GitHubLabeler\bin\Debug\netcoreapp2.0> [](start = 0, length = 55)

Is it possible to make these samples a bit less Windows specific? That is, perhaps the command to enter should be just on a line somewhere, and you elsewhere say, that it will be placed in such-and-such a directory.

TomFinley · 2018-05-22T15:01:20Z

Samples/GitHubLabeler/README.md

+This is a simple prototype application to demonstrate how to use [ML.NET](https://www.nuget.org/packages/Microsoft.ML/) APIs. The main focus is on creating, training, and using ML (Machine Learning) model that is implemented in Predictor.cs class.
+
+## Overview
+GitHubLabeler is a .NET Core console application that runs from command-line interface (CLI) and allows to:


allows to [](start = 97, length = 9)

Missing pronoun of some sort... allows you to, allows one to, other.

TomFinley · 2018-05-22T15:04:13Z

Samples/GettingStarted/Regression_TaxiFarePrediction/Program.cs

+using Microsoft.ML.Trainers;
+using Microsoft.ML.Transforms;
+using Microsoft.ML;
+using System.Threading.Tasks;


I'm not sure whether we like Microsoft before System or System before Microsoft, but we should probably pick one and make sure it's consistent.

Most developers prefer System go first and have all others alphabetically sorted afterwards.

TomFinley · 2018-05-22T15:06:43Z

Samples/GettingStarted/Regression_TaxiFarePrediction/Program.cs

+
+        private static void Evaluate(PredictionModel<TaxiTrip, TaxiTripFarePrediction> model)
+        {
+            var testData = new TextLoader<TaxiTrip>(TestDataPath, useHeader: true, separator: ",");


var testData = new TextLoader(TestDataPath, useHeader: true, separator: ","); [](start = 12, length = 87)

Something is very wrong here... the loader is part of the model, or ought to be. We shouldn't have to respecify it, any more than we have to respecify the transforms.

I think for evaluation it currently has to be specified, but this should be fixed. See #5 .

I was working on #5 and found out that Loader is actually not part of the trained model in ML.Net. I have opened another issue #216 for this blocker.

TomFinley · 2018-05-22T15:11:04Z

Samples/GitHubLabeler/GitHubLabeler/GitHubLabeler.csproj.DotSettings

@@ -0,0 +1,2 @@
+<wpf:ResourceDictionary xml:space="preserve" xmlns:x="http://schemas.microsoft.com/winfx/2006/xaml" xmlns:s="clr-namespace:System;assembly=mscorlib" xmlns:ss="urn:shemas-jetbrains-com:settings-storage-xaml" xmlns:wpf="http://schemas.microsoft.com/winfx/2006/xaml/presentation">
+	<s:String x:Key="/Default/CodeInspection/CSharpLanguageProject/LanguageLevel/@EntryValue">CSharp72</s:String></wpf:ResourceDictionary>


What's this for?

We shoudn't commit this file as this is from ReSharper ;-)

Will remove it

TomFinley · 2018-05-22T15:19:55Z

Samples/GitHubLabeler/README.md

+
+    b. To work with labels from your GitHub repository, you will need to train the model on your data. To do so, export GitHub issues from your repository into `.tsv` file with the following columns:
+    * ID - issue’s ID
+    * Area - issue’s label (named this way to avoid confusion with the Label concept in ML.NET)


’ [](start = 18, length = 1)

Should we avoid non-ascii characters unless we actually need them? I think I'd prefer a plain-old single quote here.

justinormont · 2018-05-22T15:31:47Z

Samples/GettingStarted/Regression_TaxiFarePrediction/TestTaxiTrips.cs

+            PassengerCount = 1,
+            TripDistance = 10.33f,
+            PaymentType = "CSH",
+            FareAmount = 0 // predict it. actual = 29.5


Do we allow this field to be a null or NaN? It could help the user understand this isn't an input but the output calculated value?

I don't think I can make it nullable type, getting Type mismatch exception.

eerhardt · 2018-05-22T15:45:43Z

datasets/README.md

@@ -0,0 +1,11 @@
+MICROSOFT PROVIDES THE DATASETS ON AN "AS IS" BASIS. MICROSOFT MAKES NO WARRANTIES, EXPRESS OR IMPLIED, GUARANTEES OR CONDITIONS WITH RESPECT TO YOUR USE OF THE DATASETS. TO THE EXTENT PERMITTED UNDER YOUR LOCAL LAW, MICROSOFT DISCLAIMS ALL LIABILITY FOR ANY DAMAGES OR LOSSES, INLCUDING DIRECT, CONSEQUENTIAL, SPECIAL, INDIRECT, INCIDENTAL OR PUNITIVE, RESULTING FROM YOUR USE OF THE DATASETS.


I think it would be good to put all the datasets in a single place - test/data or similar. That way they aren't duplicated unintentionally, and there is one place to see all of the datasets in the repo.

yes, was thinking about it too. I suggest moving datasets from test to the folder datasets in the root and merge them with data that is there for samples. This way it will be obvious for users where to look for it. I can do it if no objections.

There's another competing concern of having too many "things" in the root folder.

A structure I could imagine working would be:

build \ docs \ samples \ datasets \ (or just 'data'?) GettingStarted \ GitHubLabeler \ src \ Project1 \ Project2 \ test \ Test1 \ Test2 \

And all the tests changing to instead of read from test\data, they now point to samples\datasets (or samples\data).

eerhardt · 2018-05-22T15:47:44Z

...Started/BinaryClassification_SentimentAnalysis/BinaryClassification_SentimentAnalysis.csproj

+    <TargetFramework>netcoreapp2.0</TargetFramework>
+  </PropertyGroup>
+
+  <PropertyGroup Condition="'$(Configuration)|$(Platform)'=='Debug|AnyCPU'">


This PropertyGroup is incorrect, as the property below won't be set for Release builds.

If you want to set <LangVersion> to latest, it should be done unconditionally. You can move <LangVersion>latest</LangVersion> to the above PropertyGroup.

eerhardt · 2018-05-22T15:48:38Z

Samples/GitHubLabeler/GitHubLabeler/GitHubLabeler.csproj

+    <TargetFramework>netcoreapp2.0</TargetFramework>
+  </PropertyGroup>
+
+  <PropertyGroup Condition="'$(Configuration)|$(Platform)'=='Release|AnyCPU'">


Same comment here as above -

Move <LangVersion>latest</LangVersion> into the unconditional property group, and remove these two.

TomFinley · 2018-05-22T15:49:55Z

Hi @OliaG thanks so much for these examples. Let's talk about the size of the example files:

Taxi test: about 25 MB.
Taxi train: about 25 MB.
CoreFX github issue train: about 20 MB.

For machine learning these are pretty small datasets, but they're pretty big as things to commit into a git repo.

Now, we could perhaps do something like utilize Git LFS, which would be an OK remediation except that I see that this PR was done not by working out of a fork then trying to merge back into the main repo, but rather was done by establishing a branch within the main repo itself. So even if we fix it to not have these large files directly, this repo itself is somewhat tainted, since a clone or fork of the repo will contain these large files.

Do we know how to clean this up? I know it's possible, I just don't know how offhand. Maybe @eerhardt ?

Maybe development in forks a good idea, as opposed to branches in the main repo. If we look at the other PRs currently open, they're from branches in people's forks of this repo, not the repo itself. Perhaps we should clarify that somewhere. (Contribution guide, other?)

TomFinley

Hi all, sorry just waiting on changes, especially over large files. I noticed these changes were approved immediately, just want to make sure nothing bad happens.

TomFinley · 2018-05-23T03:14:39Z

datasets/imdb_labelled.txt

@@ -0,0 +1,1000 @@
+A very, very, very slow-moving, aimless movie about a distressed, drifting young man.  	0
+Not sure who was more lost - the flat characters or the audience, nearly half of whom walked out.  	0


Labeled, not labelled, FYI.

It is British VS US spelling. And it is original dataset's name, I did not change it. But I can remove extra "l" :)

When in Rome ...

Oh really? Hmmm. I never knew that. How strange. I guess if that's the original name we should keep it.

* added ceiling to mathutils and have parquetloader default to sequential reading instead of throwing * factor out sequence creation and add check for overflow in MathUtils

…ph variable outputs (#152) * Adding support for training metrics in PipelineSweeperMacro and needed support files. Also includes new output information in PipelineSweeperMacro output graph to make consumption of returned pipelines easier. * Changed where XML comment was placed. * Added more tests (uncommented and fixed) for auto inference. Changed magic number in AutoMlUtils to be mix double value, per review comment. * Added another test, TestPipelineSweeperMacroNoTransforms. * Updated tests to include warning disabling (following Zeeshan S's example) to get build working. * Changes to checks in AutoMlUtils (more correct usage of them). * Fixing issue with ExceptParam using value and not name of parameters. * Fixing errors on use of ExceptParam.

terrajobst · 2018-05-23T16:38:10Z

Samples/GettingStarted/BinaryClassification_SentimentAnalysis/SentimentData.cs

+{
+    public class SentimentData
+    {
+        [Column("0")] public string SentimentText;


I suggest putting the attributes on separate lines as that's more in line with default formatting.

terrajobst · 2018-05-23T16:39:41Z

Samples/GettingStarted/MulticlassClassification_Iris/Program.cs

+    {
+        private static string AppPath => Path.GetDirectoryName(Environment.GetCommandLineArgs()[0]);
+        private static string TrainDataPath => Path.Combine(AppPath, @"..\..\..\..\datasets\", "iris_train.txt");
+        private static string TestDataPath => Path.Combine(AppPath, @"..\..\..\..\datasets\", "iris_test.txt");


Don't use backslashes as this doesn't work cross-platform. You should do:

Path.Combine(AppPath, "..", "..", "..", "..", "datasets", "iris_train.txt");

terrajobst · 2018-05-23T16:40:53Z

Samples/GettingStarted/Regression_TaxiFarePrediction/Program.cs

+using Microsoft.ML.Trainers;
+using Microsoft.ML.Transforms;
+using Microsoft.ML;
+using System.Threading.Tasks;


Most developers prefer System go first and have all others alphabetically sorted afterwards.

terrajobst · 2018-05-23T16:41:44Z

Samples/GitHubLabeler/.gitignore

@@ -0,0 +1,288 @@
+## Ignore Visual Studio temporary files, build results, and
+## files generated by popular Visual Studio add-ons.
+##


Agreed. We should only use the one from the root.

terrajobst · 2018-05-23T16:43:35Z

Samples/GitHubLabeler/GitHubLabeler/App.config

+<?xml version="1.0" encoding="utf-8" ?>
+<configuration>
+  <appSettings>
+    <!--TODO: Please enter your own credentials here. -->


I don't think there is a great solution for this. In general, the recommendation is to avoid using credentials from code and instead use certificates, user-based authentication, or a service like Azure KeyVault. Unfortunately this doesn't work for samples very well...

terrajobst · 2018-05-23T16:44:21Z

Samples/GitHubLabeler/GitHubLabeler/GitHubLabeler.csproj.DotSettings

@@ -0,0 +1,2 @@
+<wpf:ResourceDictionary xml:space="preserve" xmlns:x="http://schemas.microsoft.com/winfx/2006/xaml" xmlns:s="clr-namespace:System;assembly=mscorlib" xmlns:ss="urn:shemas-jetbrains-com:settings-storage-xaml" xmlns:wpf="http://schemas.microsoft.com/winfx/2006/xaml/presentation">
+	<s:String x:Key="/Default/CodeInspection/CSharpLanguageProject/LanguageLevel/@EntryValue">CSharp72</s:String></wpf:ResourceDictionary>


We shoudn't commit this file as this is from ReSharper ;-)

terrajobst · 2018-05-23T16:44:59Z

Samples/GitHubLabeler/README.md

+
+    b. To work with labels from your GitHub repository, you will need to train the model on your data. To do so, export GitHub issues from your repository into `.tsv` file with the following columns:
+    * ID - issue’s ID
+    * Area - issue’s label (named this way to avoid confusion with the Label concept in ML.NET)


* Reduce number of hash bits in stratification column and add a unit test. * Address PR comments.

* replace housing uci dataset to wine quality

Moved datasets from tests to examples Removed model files

- Iris sample - Sentiment Analysis sample - TaxiFare sample -GitHub issues classification sample

Moved datasets from tests to examples Removed model files

…nto samples

KrzysztofCwalina · 2018-05-24T15:10:48Z

This is a very nice PR, but shouldn't we keep all samples/examples in the same place? So far, we have been collecting such samples in tests/Microsoft/ML/Scenarios.

OliaG · 2018-05-24T16:52:22Z

@KrzysztofCwalina regarding keep all samples/example in the same place, here is my understanding.

test/Microsoft/ML/Scenarios are integration tests the primary goal of which is to test our code end-to-end. We can run them to make sure there are no breaking changes introduced by new functionality.

The examples (this PR) are a way of showcasing our functionality to users. They are not tests but independent projects, sometimes applications containing additional logic (like posting requests to GitHub). The examples have to be discoverable for users (like root folder “examples”, it will be hard for users to find them here: test/Microsoft/ML/Scenarios) and F5-ble, users should be able do download just one example and successfully run it. You can’t do it with integration tests.

So for me they belong to different areas and that’s why I would keep tests in tests, and this samples in examples folder. It is a good idea though to mention these tests in docs or readme, so users can find additional use cases.

TomFinley · 2018-05-24T16:57:29Z

docs/code/IdvFileFormat.md

@@ -0,0 +1,191 @@
+# IDV File Format
+
+This document describes ML.NET's Binary dataview file format, version 1.1.1.5


This looks familiar... was this intended to be part of this PR? Was there a bad merge somewhere?

* removed dummy command * sanitize label and fix template * fix tests

@daholste

…ature branch (#3324) * Initial commit * ci test build * forgot to save this one file * Debug-Intrinsics isn't a valid config, trying windows-x64 * disabled tests for now * disable tests attempt 2 * initial code push, no history, test project not in the build so is the internal client * battling with warn as err * test build * test change * make params for MLContext data extensions match ML.NET default names and values; update gitignore; nit rev for Benchmarking.cs (#5) * Create README.md (#2) * API folder changes (#6) * comment out fast forest trainer, per discussion on ML.NET open issue #1983, for now, to run E2E w/o exceptions (#7) * Make validation data param mandatory; remove GetFirstPipeline sample (#10) * Make validation data param mandatory; remove GetFirstPipeline sample * remove deprecated todo * Create ISSUE_TEMPLATE.md & PULL_REQUEST_TEMPLATE.md (#12) * Create ISSUE_TEMPLATE.md * Create PULL_REQUEST_TEMPLATE.md * NestedObject For pipeline (#14) * add estimator extensions / catalog; add conversion from external to internal pipeline; transform clean-up; add back in test proj and fix build; refactor trainer ext name mappings (#15) * Make validation data param mandatory; remove GetFirstPipeline sample * remove deprecated todo * add estimator extensions / catalog; add ability to go from external to internal pipeline; a lot of transform clean-up; add back in test proj and get it building; refactor trainer ext name mappings * corrected the typo in readme (#16) * make GetNextPipeline API w/ public Pipeline method on PipelineSuggester; write GetNextPipeline API test; fix public Pipeline object serialization; fix header inferencing bug; write test utils for fetching datasets (#18) * get next pipeline API rev -- refactor API to consume column dimensions, purpose, type, and name instead of available trainers & transforms (#19) * mark get next pipeline test as ignore for now (#20) * fix dataview take util bug, add dataview skip util, add some UTs to increase code coverage (#21) * fix dataview take util bug, add dataview skip util, add some UTs to increase code coverage * add accuracy threshold on AutoFit test * add null check to best pipeline on autofit result * unit test additions (including user input validation testing); dead code removal for code coverage (including KDO & associated utils); misc fixes & revs (#22) * add trainer extension tests, & misc fixes (#23) * add estimator extension tests (#24) * add conversions tests (#25) * fix multiclass runs & add multiclass autofit UT (#27) * add basic autofit regression test (#28) * fix categorical transform bug (sometimes categorical features weren't concatenated to final features); add UT transforms; add PipelineNode equality & tests to serve as AutoML testing infra * add example to readme (#26) * add lightgbm args as nested properties (#33) * fix bug where if one pipeline hyperparam optimization converges, run terminates (#36) * add open-source headers to files; other nit clean-ups along the way (#35) * Ungroup Columns in Column Inference (#40) * Added sequential grouping of columns * added ungrouping of column option * reverted the file * Misc fixes (#39) * misc fixes -- fix bug where SMAC returning already-seen values; fix param encoding return bug in pipeline object model; nit clean-up AutoFit; return in pipeline suggester when sweeper has no next proposal; null ref fix in public object model pipeline suggester * fix in BuildPipelineNodePropsLightGbm test, fix / use correct 'newTrainer' variable in PipelneSuggester * SMAC perf improvement * Removing the nuget.config and have build.props mention the nuget package sources. (#38) * Added sequential grouping of columns * removed nuget.config and have only props mentions the nuget sources * reverted the file * transform inferencing concat / ignore fixes (#41) * make pipeline object model & other public classes internal (#43) * handle SMAC exception when fewer trees were trained than requested (#44) * Throw error on incorrect Label name in InferColumns API (#47) * Added sequential grouping of columns * reverted the file * addded infer columns label name checking * added column detection error * removed unsed usings * added quotes * replace Where with Any clause * replace Where with Any clause * Set Nullable Auto params to null values (#50) * Added sequential grouping of columns * reverted the file * added auto params as null * change to the update fields method * First public api propsal (#52) * Includes following 1) Final proposal for 0.1 public API surface 2) Prefeaturization 3) Splitting train data into train and validate when validation data is null 4) Providing end to end samples one each for regression, binaryclassification and multiclass classification * Incorporating code review feedbacks * Revert "Set Nullable Auto params to null values" (#53) * Revert "First public api propsal (#52)" This reverts commit e4a64cf. * Revert "Set Nullable Auto params to null values (#50)" This reverts commit 41c663c. * AutoFit return type is now an IEnumerable (#55) AutoFit returns is now an IEnumerable - this enables many good things Implementing variety of early stopping criteria (See sample) Early discard of models that are no good. This improves memory usage efficiency. (See sample) No need to implement a callback to get results back Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample). Also templatized the return type for better type safety through out the code. * misc fixes & test additions, towards 0.1 release (#56) * Enable UnitTests on build server (#57) * 1) Making trainer name public (#62) 2) Fixing up samples to reflect it * Initial version of CLI tool for mlnet (#61) * added global tool initial project * removed unneccesary files, renamed files * refactoring and added base abstract classes for trainer generator * removed unused class * Added classes for transforms * added transform generate dummy classes * more refactoring, added first transform * more refactoring and added classes * changed the project structure * restructing added options class * sln changes * refactored options to different class: * added more logic for code generation of class * misc changes * reverted file * added commandline api package * reverted sample * added new command line api parser * added normalization of column names * Added command defaults and error message * implementation of all trainers * changed auto to null * added all transform generators * added error handling when args is empty and minor changes due to change in AutoML api names * changed the name of param * added new command line options and restructuring code * renamed proj file and added solution * Added code to generate usings, Fixed few bugs in the code * added validation to the command line options * changed project name * Bug fixes due to API change in AutoML * changed directory structure * added test framework and basic tests * added more tests * added improvements to template and error handling * renamed the estimator name * fixed test case * added comments * added headers * changed namespace and removed unneccesary properties from project * Revert "changed namespace and removed unneccesary properties from project" This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f. * fixed test cases and renamed namespaces * cleaned up proj file * added folder structure * added symbols/tokens for strings * added more tests * review comments * modified test cases * review comments * change in the exception message * normalized line endings * made method private static * simplified range building /optimization * minor fix * added header * added static methods in command where necessary * nit picks * made few methods static * review comments * nitpick * remove line pragmas * fix test case * Use better AutiFit overload and ignore Multiclass (#64) * Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65) * Added sequential grouping of columns * reverted the file * upgrade to v .10 and refactoring * added null check * fixed unit tests * review comments * removed the settings change * added regions * fixed unit tests * Upgrade ML.NET package to 0.10.0 (#70) * Change in template to accomodate new API of TextLoader (#72) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * Enable gated check for mlnet.tests (#79) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * added run-tests.proj and referred it in build.proj * CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83) * Added sequential grouping of columns * reverted the file * bug fixes, more logic to templates to support cross-validate * formatting and fix type in consolehelper * Added logic in templates * revert settings * benchmarking related changes (#63) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * fix fast forest learner (don't sweep over learning rate) (#88) * Made changes to Have non-calibrated scoring for binary classifiers (#86) * Added sequential grouping of columns * reverted the file * added calibration workaround * removed print probability * reverted settings * rev ColumnInference API: can take label index; rev output object types; add tests (#89) * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * publish nuget (#101) * use dotnet-internal-temp agent for internal build * use dotnet-internal feed * Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95) * Added sequential grouping of columns * reverted the file * fix usings for type convert * added transforms tests * review comments * When generating usings choose only distinct usings directives (#94) * Added sequential grouping of columns * reverted the file * Added code to have unique strings * refactoring * minor fix * minor fix * Autofit overloads + cancellation + progress callbacks 1) Introduce AutoFit overloads (basic and advanced) 2) AutoFit Cancellation 3) AutoFit progress callbacks * Default the kfolds to value 5 in CLI generated code (#115) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * remove file * added kfold param and defaulted to value * changed type * added for regression * Remove extra ; from generated code (#114) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * removed extra ; from generated code * removed file * fix unit tests * TimeoutInSeconds (#116) Specifying timeout in seconds instead of minutes * Added more command line args implementation to CLI tool and refactoring (#110) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * added git status * reverted change * added codegen options and refactoring * minor fixes' * renamed params, minor refactoring * added tests for commandline and refactoring * removed file * added back the test case * minor fixes * Update src/mlnet.Test/CommandLineTests.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * review comments * capitalize the first character * changed the name of test case * remove unused directives * Fail gracefully if unable to instantiate data view with swept parameters (#125) * gracefully fail if fail to parse a datai * rev * validate AutoFit 'Features' column must be of type R4 (#132) * Samples: exceptions / nits (#124) * Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121) * addded logging and helper methods * fixing code after merge * added resx files, added logger framework, added logging messages * added new options * added spacing * minor fixes * change command description * rename option, add headers, include new param in test * formatted * build fix * changed option name * Added NlogConfig file * added back config package * fix tests * added correct validation check (#137) * Use CreateTextLoader<T>(..) instead of CreateTextLoader(..) (#138) * added support to loaddata by class in the generated code * fix tests * changed CreateTextLoader to ReadFromTextFile method. (#140) * changed textloader to readfromtextfile method * formatting * exception fixes (#136) * infer purpose of hidden columns as 'ignore' (#142) * Added approval tests and bunch of refactoring of code and normalizing namespaces (#148) * changed textloader to readfromtextfile method * formatting * added approval tests and refactoring of code * removed few comments * API 2.0 skeleton (#149) Incorporating API review feedback * The CV code should come before the training when there is no test dataset in generated code (#151) * reorder cv code * build fix * fixed structure * Format the generated code + bunch of misc tasks (#152) * added formatting and minor changes for reordering cv * fixing the template * minor changes * formatting changes * fixed approval test * removed unused nuget * added missing value replacing * added test for new transform * fix test * Update src/mlnet/Templates/Console/MLCodeGen.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Sanitize the column names in CLI (#162) * added sanitization layer in CLI * fix test * changed exception.StackTrace to exception.ToString() * fix package name (#168) * Rev public API (#163) * Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153) * Fix minor version for the repository + remove Nlog config package (#171) * changed the minor version * removed the nlog config package * Added new test to columninfo and fixing up API (#178) * Make optimizing metric customizable and add trainer whitelist functionality (#172) * API rev (#181) * propagate root MLContext thru AutoML (instead of creating our own) (#182) * Enabling new command line args (#183) * fix package name * initial commit * added more commandline args * fixed tests * added headers * fix tests * fix test * rename 'AutoFitter' to 'Experiment' (#169) * added tests (#187) * rev InferColumns to accept ColumnInfo input param (#186) * Implement argument --has-header and change usage of dataset (#194) * added has header and fixed dataset and train dataset * fix tests * removed dummy command (#195) * Fix bug for regression and sanitize input label from user (#198) * removed dummy command * sanitize label and fix template * fix tests * Do not generate code concatenating columns when the dataset has a single feature column (#191) * Include some missed logging in the generated code. (#199) * added logging messages for generated code * added log messages * deleted file * cleaning up proj files (#185) * removed platform target * removed platform target * Some spaces and extra lines + bug in output path (#204) * nit picks * nit picks * fix test * accept label from user input and provide in generated code (#205) * Rev handling of weight / label columns (#203) * migrate to private ML.NET nuget for latest bug fixes (#131) * fix multiclass with nonstandard label (#207) * Multiclass nondefault label test (#208) * printing escaped chars + bug (#212) * delete unused internal samples (#211) * fix SMAC bug that causes multiclass sample to infinite loop (#209) * Rev user input validation for new API (#210) * added console message for exit and nit picks (#215) * exit when exception encountered (#216) * Seal API classes (and make EnableCaching internal) (#217) * Suggested sample nits (feel free to ask for any of these to be reverted) (#219) * User input column type validation (#218) * upgrade commandline and renaming (#221) * upgrade commandline and renaming * renaming fields * Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225) * CLI argument descriptions updated (#224) * CLI argument descriptions updated * No version in .csproj * added flag to disable training code (#227) * Exit if perfect model produced (#220) * removed header (#228) * removed header * added auto generated header * removed console read key (#229) * Fix model path in generated file (#230) * removed console read key * fix model path * fix test * reorder samples (#231) * remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233) * Null reference exception fix for finding best model when some runs have failed (#239) * samples fixes (#238) * fix for defaulting Averaged Perceptron # of iterations to 10 (#237) * Bug bash feedback Feb 27. API changes and sample changes (#240) * Bug bash feedback Feb 27. API changes Sample changes Exception fix * Samples / API rev from 2/27 bug bash feedback (#242) * changed the directory structure for generated project (#243) * changed the directory structure for generated project * changed test * upgraded commandline package * Fix test file locations on OSX (#235) * fix test file locations on OSX * changing to Path.Combine() * Additional Path.Combine() * Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt * Additional Path.Combine() * add back in double comparison fix * remove metrics agent NaN returns * test fix * test format fix * mock out path Thanks to @daholste for additional fixes! * upgrade to latest ML.NET public surface (#246) * Upgrade to ML.NET 0.11 (#247) * initial changes * fix lightgbm * changed normalize method * added tests * fix tests * fix test * Private preview final API changes (#250) * .NET framework design guidelines applied to public surface * WhitelistedTrainers -> Trainers * Add estimator to public API iteration result (#248) * LightGBM pipeline serialization fix (#251) * Change order that we search for TextLoader's parameters (#256) * CLI IFileInfo null exception fix (#254) * Averaged Perceptron pipeline serialization fix (#257) * Upgrade command-line-api and default folder name change (#258) * change in defautl folderName * upgrade command line * Update src/mlnet/Program.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * eliminate IFileInfo from CLI (#260) * Rev samples towards private preview; ignored columns fix (#259) * remove unused methods in consolehelper and nit picks in generated code (#261) * nit picks * change in console helper * fix tests * add space * fix tests * added nuget sources in generated csproj (#262) * added nuget sources in csproj * changed the structure in generated code * space * upgrade to mlnet 0.11 (#263) * Formatting CLI metrics (#264) Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits. * Add implementation of non -ova multi class trainers code gen (#267) * added non ova multi class learners * added tests * test cases * Add caching (#249) * AdvancedExperimentSettings sample nits (#265) * Add sampling key column (#268) * Initial work for multi-class classification support for CLI (#226) * Initial work for multi-class classification support for CLI * String updates * more strings * Whitelist non-OVA multi-class learners * Refactor the orchestration of AutoML calls (#272) * Do not auto-group columns with suggested purpose = 'Ignore' (#273) * Fix: during type inferencing, parse whitespace strings as NaN (#271) * Printing additional metrics in CLI for binary classification (#274) * Printing additional metrics in CLI for binary classification * Update src/mlnet/Utilities/ConsolePrinter.cs * Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269) * Print failed iterations in CLI (#275) * change the type to float from double (#277) * cache arg implementation in CLI (#280) * cache implementation * corrected the null case * added tests for all cases * Remove duplicate value-to-key mapping transform for multiclass string labels (#283) * Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286) * Implement ignore columns command line arg (#290) * normalize line endings * added --ignore-columns * null checks * unit tests * Print winning iteration and runtime in CLI (#288) * Print best metric and runtime * Print best metric and runtime * Line endings in AutoMLEngine.cs * Rename time column to duration to match Python SDK * Revert to MicroAccuracy and MacroAccuracy spellings * Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts * Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts * missed some files * Fix merge conflict * Update AutoMLEngine.cs * Add MacOS & Linux to CI; MacOS & Linux test fixes (#293) * MicroAccuracy as default for multi-class (#295) Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy. * Null exception for ignorecolumns in CLI (#294) * Null exception for ignorecolumns in CLI * Check if ignore-columns array has values (as the default is now a empty array) * Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296) * removed sln (#297) * Caching enabling in code gen part -2 (#298) * add * added caching codegen * support comma separated values for --ignore-columns (#300) * default initialization for ignore columns (#302) * default initialization * adde null check * Codegen for multiclass non-ova (#303) * changes to template * multicalss codegen * test cases * fix test cases * Generated Project new structure. (#305) * added new templates * writing files to disck * change path * added new templates * misisng braces * fix bugs * format code * added util methods for solution file creation and addition of projects to it * added extra packages to project files * new tests * added correct path for sln * build fix * fix build * include using system in prediction class (#307) * added using * fix test * Random number generator is not thread safe (#310) * Random number generator is not thread safe * Another local random generator * Missed a few references * Referncing AutoMlUtils.random instead of a local RNG * More refs to mail RNG; remove Float as per #1669 * Missed Random.cs * Fix multiclass code gen (#314) * compile error in codegen * removes scores printing * fix bugs * fix test * Fix compile error in codegen project (#319) * removed redundant code * fix test case * Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317) * Ova Multi class codegen support (#321) * dummy * multiova implementation * fix tests * remove inclusion list * fix tests and console helper * Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322) * Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination * test fixes * Console helper bug in generated code for multiclass (#323) * fix * fix test * looping perlogclass * fix test * Initial version of Progress bar impl and CLI UI experience (#325) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * Setting model directory to temp directory (#327) * Suggested changes to progress bar (#335) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * Rev Samples (#334) * Telemetry2 (#333) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * CLI telemetry implementation * Telemetry implementation * delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value * add headers, remove comments * one more header missing * Fix progress bar in linux/osx (#336) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * change from task to thread * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Mem leak fix (#328) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * there is still investigation to be done but this fix works and solves memory leak problems * minor refactor * Upgrade ML.NET package (#343) * Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287) * restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344) * Polishing the CLI UI part-1 (#338) * formatting of pbar message * Polishing the UI * optimization * rename variable * Update src/mlnet/AutoML/AutoMLEngine.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * new message * changed hhtp to https * added iteration num + 1 * change string name and add color to artifacts * change the message * build errors * added null checks * added exception messsages to log file * added exception messsages to log file * CLI ML.NET version upgrade (#345) * Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346) * CLI -- consume logs from AutoML SDK (#349) * Rename RunDetails --> RunDetail (#350) * command line api upgrade and progress bar rendering bug (#366) * added fix for all platforms progress bar * upgrade nuget * removed args from writeline * change in the version (#368) * fix few bugs in progressbar and verbosity (#374) * fix few bugs in progressbar and verbosity * removed unused name space * Fix for folders with space in it while generating project (#376) * support for folders with spaces * added support for paths with space * revert file * change name of var * remove spaces * SMAC fix for minimizing metrics (#363) * Formatting Regression metrics and progress bar display days. (#379) * added progress bar day display and fix regression metrics * fix formatting * added total time * formatted total time * change command name and add pbar message (#380) * change command name and add pbar message * fix tests * added aliases * duplicate alias * added another alias for task * UI missing features (#382) * added formatting changes * added accuracy specifically * downgrade the codepages (#384) * Change in project structure (#385) * initial changes * Change in project structure * correcting test * change variable name * fix tests * fix tests * fix more tests * fix codegen errors * adde log file message * changed name of args * change variable names * fix test * FileSizeBuckets in correct units (#387) * Minor telemetry change to log in correct units and make our life easier in the future * Use Ceiling instead of Round * changed order (#388) * prep work to transfer to ml.net (#389) * move test projects to top level test subdir * rename some projects to make naming consistent and make it build again * fix test project refs * Add AutoML components to build, fix issues related to that so it builds

* removed dummy command * sanitize label and fix template * fix tests

* Fixed build errors resulting from upgrade to VS2019 compilers * Added additional message describing the previous fix * Syncing upstream fork (#10) * Throw error on incorrect Label name in InferColumns API (#47) * Added sequential grouping of columns * reverted the file * addded infer columns label name checking * added column detection error * removed unsed usings * added quotes * replace Where with Any clause * replace Where with Any clause * Set Nullable Auto params to null values (#50) * Added sequential grouping of columns * reverted the file * added auto params as null * change to the update fields method * First public api propsal (#52) * Includes following 1) Final proposal for 0.1 public API surface 2) Prefeaturization 3) Splitting train data into train and validate when validation data is null 4) Providing end to end samples one each for regression, binaryclassification and multiclass classification * Incorporating code review feedbacks * Revert "Set Nullable Auto params to null values" (#53) * Revert "First public api propsal (#52)" This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7. * Revert "Set Nullable Auto params to null values (#50)" This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d. * AutoFit return type is now an IEnumerable (#55) AutoFit returns is now an IEnumerable - this enables many good things Implementing variety of early stopping criteria (See sample) Early discard of models that are no good. This improves memory usage efficiency. (See sample) No need to implement a callback to get results back Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample). Also templatized the return type for better type safety through out the code. * misc fixes & test additions, towards 0.1 release (#56) * Enable UnitTests on build server (#57) * 1) Making trainer name public (#62) 2) Fixing up samples to reflect it * Initial version of CLI tool for mlnet (#61) * added global tool initial project * removed unneccesary files, renamed files * refactoring and added base abstract classes for trainer generator * removed unused class * Added classes for transforms * added transform generate dummy classes * more refactoring, added first transform * more refactoring and added classes * changed the project structure * restructing added options class * sln changes * refactored options to different class: * added more logic for code generation of class * misc changes * reverted file * added commandline api package * reverted sample * added new command line api parser * added normalization of column names * Added command defaults and error message * implementation of all trainers * changed auto to null * added all transform generators * added error handling when args is empty and minor changes due to change in AutoML api names * changed the name of param * added new command line options and restructuring code * renamed proj file and added solution * Added code to generate usings, Fixed few bugs in the code * added validation to the command line options * changed project name * Bug fixes due to API change in AutoML * changed directory structure * added test framework and basic tests * added more tests * added improvements to template and error handling * renamed the estimator name * fixed test case * added comments * added headers * changed namespace and removed unneccesary properties from project * Revert "changed namespace and removed unneccesary properties from project" This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f. * fixed test cases and renamed namespaces * cleaned up proj file * added folder structure * added symbols/tokens for strings * added more tests * review comments * modified test cases * review comments * change in the exception message * normalized line endings * made method private static * simplified range building /optimization * minor fix * added header * added static methods in command where necessary * nit picks * made few methods static * review comments * nitpick * remove line pragmas * fix test case * Use better AutiFit overload and ignore Multiclass (#64) * Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65) * Added sequential grouping of columns * reverted the file * upgrade to v .10 and refactoring * added null check * fixed unit tests * review comments * removed the settings change * added regions * fixed unit tests * Upgrade ML.NET package to 0.10.0 (#70) * Change in template to accomodate new API of TextLoader (#72) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * Enable gated check for mlnet.tests (#79) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * added run-tests.proj and referred it in build.proj * CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83) * Added sequential grouping of columns * reverted the file * bug fixes, more logic to templates to support cross-validate * formatting and fix type in consolehelper * Added logic in templates * revert settings * benchmarking related changes (#63) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * fix fast forest learner (don't sweep over learning rate) (#88) * Made changes to Have non-calibrated scoring for binary classifiers (#86) * Added sequential grouping of columns * reverted the file * added calibration workaround * removed print probability * reverted settings * rev ColumnInference API: can take label index; rev output object types; add tests (#89) * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * publish nuget (#101) * use dotnet-internal-temp agent for internal build * use dotnet-internal feed * Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95) * Added sequential grouping of columns * reverted the file * fix usings for type convert * added transforms tests * review comments * When generating usings choose only distinct usings directives (#94) * Added sequential grouping of columns * reverted the file * Added code to have unique strings * refactoring * minor fix * minor fix * Autofit overloads + cancellation + progress callbacks 1) Introduce AutoFit overloads (basic and advanced) 2) AutoFit Cancellation 3) AutoFit progress callbacks * Default the kfolds to value 5 in CLI generated code (#115) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * remove file * added kfold param and defaulted to value * changed type * added for regression * Remove extra ; from generated code (#114) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * removed extra ; from generated code * removed file * fix unit tests * TimeoutInSeconds (#116) Specifying timeout in seconds instead of minutes * Added more command line args implementation to CLI tool and refactoring (#110) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * added git status * reverted change * added codegen options and refactoring * minor fixes' * renamed params, minor refactoring * added tests for commandline and refactoring * removed file * added back the test case * minor fixes * Update src/mlnet.Test/CommandLineTests.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * review comments * capitalize the first character * changed the name of test case * remove unused directives * Fail gracefully if unable to instantiate data view with swept parameters (#125) * gracefully fail if fail to parse a datai * rev * validate AutoFit 'Features' column must be of type R4 (#132) * Samples: exceptions / nits (#124) * Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121) * addded logging and helper methods * fixing code after merge * added resx files, added logger framework, added logging messages * added new options * added spacing * minor fixes * change command description * rename option, add headers, include new param in test * formatted * build fix * changed option name * Added NlogConfig file * added back config package * fix tests * added correct validation check (#137) * Use CreateTextLoader<T>(..) instead of CreateTextLoader(..) (#138) * added support to loaddata by class in the generated code * fix tests * changed CreateTextLoader to ReadFromTextFile method. (#140) * changed textloader to readfromtextfile method * formatting * exception fixes (#136) * infer purpose of hidden columns as 'ignore' (#142) * Added approval tests and bunch of refactoring of code and normalizing namespaces (#148) * changed textloader to readfromtextfile method * formatting * added approval tests and refactoring of code * removed few comments * API 2.0 skeleton (#149) Incorporating API review feedback * The CV code should come before the training when there is no test dataset in generated code (#151) * reorder cv code * build fix * fixed structure * Format the generated code + bunch of misc tasks (#152) * added formatting and minor changes for reordering cv * fixing the template * minor changes * formatting changes * fixed approval test * removed unused nuget * added missing value replacing * added test for new transform * fix test * Update src/mlnet/Templates/Console/MLCodeGen.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Sanitize the column names in CLI (#162) * added sanitization layer in CLI * fix test * changed exception.StackTrace to exception.ToString() * fix package name (#168) * Rev public API (#163) * Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153) * Fix minor version for the repository + remove Nlog config package (#171) * changed the minor version * removed the nlog config package * Added new test to columninfo and fixing up API (#178) * Make optimizing metric customizable and add trainer whitelist functionality (#172) * API rev (#181) * propagate root MLContext thru AutoML (instead of creating our own) (#182) * Enabling new command line args (#183) * fix package name * initial commit * added more commandline args * fixed tests * added headers * fix tests * fix test * rename 'AutoFitter' to 'Experiment' (#169) * added tests (#187) * rev InferColumns to accept ColumnInfo input param (#186) * Implement argument --has-header and change usage of dataset (#194) * added has header and fixed dataset and train dataset * fix tests * removed dummy command (#195) * Fix bug for regression and sanitize input label from user (#198) * removed dummy command * sanitize label and fix template * fix tests * Do not generate code concatenating columns when the dataset has a single feature column (#191) * Include some missed logging in the generated code. (#199) * added logging messages for generated code * added log messages * deleted file * cleaning up proj files (#185) * removed platform target * removed platform target * Some spaces and extra lines + bug in output path (#204) * nit picks * nit picks * fix test * accept label from user input and provide in generated code (#205) * Rev handling of weight / label columns (#203) * migrate to private ML.NET nuget for latest bug fixes (#131) * fix multiclass with nonstandard label (#207) * Multiclass nondefault label test (#208) * printing escaped chars + bug (#212) * delete unused internal samples (#211) * fix SMAC bug that causes multiclass sample to infinite loop (#209) * Rev user input validation for new API (#210) * added console message for exit and nit picks (#215) * exit when exception encountered (#216) * Seal API classes (and make EnableCaching internal) (#217) * Suggested sample nits (feel free to ask for any of these to be reverted) (#219) * User input column type validation (#218) * upgrade commandline and renaming (#221) * upgrade commandline and renaming * renaming fields * Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225) * CLI argument descriptions updated (#224) * CLI argument descriptions updated * No version in .csproj * added flag to disable training code (#227) * Exit if perfect model produced (#220) * removed header (#228) * removed header * added auto generated header * removed console read key (#229) * Fix model path in generated file (#230) * removed console read key * fix model path * fix test * reorder samples (#231) * remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233) * Null reference exception fix for finding best model when some runs have failed (#239) * samples fixes (#238) * fix for defaulting Averaged Perceptron # of iterations to 10 (#237) * Bug bash feedback Feb 27. API changes and sample changes (#240) * Bug bash feedback Feb 27. API changes Sample changes Exception fix * Samples / API rev from 2/27 bug bash feedback (#242) * changed the directory structure for generated project (#243) * changed the directory structure for generated project * changed test * upgraded commandline package * Fix test file locations on OSX (#235) * fix test file locations on OSX * changing to Path.Combine() * Additional Path.Combine() * Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt * Additional Path.Combine() * add back in double comparison fix * remove metrics agent NaN returns * test fix * test format fix * mock out path Thanks to @daholste for additional fixes! * upgrade to latest ML.NET public surface (#246) * Upgrade to ML.NET 0.11 (#247) * initial changes * fix lightgbm * changed normalize method * added tests * fix tests * fix test * Private preview final API changes (#250) * .NET framework design guidelines applied to public surface * WhitelistedTrainers -> Trainers * Add estimator to public API iteration result (#248) * LightGBM pipeline serialization fix (#251) * Change order that we search for TextLoader's parameters (#256) * CLI IFileInfo null exception fix (#254) * Averaged Perceptron pipeline serialization fix (#257) * Upgrade command-line-api and default folder name change (#258) * change in defautl folderName * upgrade command line * Update src/mlnet/Program.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * eliminate IFileInfo from CLI (#260) * Rev samples towards private preview; ignored columns fix (#259) * remove unused methods in consolehelper and nit picks in generated code (#261) * nit picks * change in console helper * fix tests * add space * fix tests * added nuget sources in generated csproj (#262) * added nuget sources in csproj * changed the structure in generated code * space * upgrade to mlnet 0.11 (#263) * Formatting CLI metrics (#264) Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits. * Add implementation of non -ova multi class trainers code gen (#267) * added non ova multi class learners * added tests * test cases * Add caching (#249) * AdvancedExperimentSettings sample nits (#265) * Add sampling key column (#268) * Initial work for multi-class classification support for CLI (#226) * Initial work for multi-class classification support for CLI * String updates * more strings * Whitelist non-OVA multi-class learners * Refactor the orchestration of AutoML calls (#272) * Do not auto-group columns with suggested purpose = 'Ignore' (#273) * Fix: during type inferencing, parse whitespace strings as NaN (#271) * Printing additional metrics in CLI for binary classification (#274) * Printing additional metrics in CLI for binary classification * Update src/mlnet/Utilities/ConsolePrinter.cs * Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269) * Print failed iterations in CLI (#275) * change the type to float from double (#277) * cache arg implementation in CLI (#280) * cache implementation * corrected the null case * added tests for all cases * Remove duplicate value-to-key mapping transform for multiclass string labels (#283) * Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286) * Implement ignore columns command line arg (#290) * normalize line endings * added --ignore-columns * null checks * unit tests * Print winning iteration and runtime in CLI (#288) * Print best metric and runtime * Print best metric and runtime * Line endings in AutoMLEngine.cs * Rename time column to duration to match Python SDK * Revert to MicroAccuracy and MacroAccuracy spellings * Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts * Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts * missed some files * Fix merge conflict * Update AutoMLEngine.cs * Add MacOS & Linux to CI; MacOS & Linux test fixes (#293) * MicroAccuracy as default for multi-class (#295) Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy. * Null exception for ignorecolumns in CLI (#294) * Null exception for ignorecolumns in CLI * Check if ignore-columns array has values (as the default is now a empty array) * Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296) * removed sln (#297) * Caching enabling in code gen part -2 (#298) * add * added caching codegen * support comma separated values for --ignore-columns (#300) * default initialization for ignore columns (#302) * default initialization * adde null check * Codegen for multiclass non-ova (#303) * changes to template * multicalss codegen * test cases * fix test cases * Generated Project new structure. (#305) * added new templates * writing files to disck * change path * added new templates * misisng braces * fix bugs * format code * added util methods for solution file creation and addition of projects to it * added extra packages to project files * new tests * added correct path for sln * build fix * fix build * include using system in prediction class (#307) * added using * fix test * Random number generator is not thread safe (#310) * Random number generator is not thread safe * Another local random generator * Missed a few references * Referncing AutoMlUtils.random instead of a local RNG * More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669 * Missed Random.cs * Fix multiclass code gen (#314) * compile error in codegen * removes scores printing * fix bugs * fix test * Fix compile error in codegen project (#319) * removed redundant code * fix test case * Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317) * Ova Multi class codegen support (#321) * dummy * multiova implementation * fix tests * remove inclusion list * fix tests and console helper * Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322) * Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination * test fixes * Console helper bug in generated code for multiclass (#323) * fix * fix test * looping perlogclass * fix test * Initial version of Progress bar impl and CLI UI experience (#325) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * Setting model directory to temp directory (#327) * Suggested changes to progress bar (#335) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * Rev Samples (#334) * Telemetry2 (#333) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * CLI telemetry implementation * Telemetry implementation * delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value * add headers, remove comments * one more header missing * Fix progress bar in linux/osx (#336) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * change from task to thread * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Mem leak fix (#328) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * there is still investigation to be done but this fix works and solves memory leak problems * minor refactor * Upgrade ML.NET package (#343) * Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287) * restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344) * Polishing the CLI UI part-1 (#338) * formatting of pbar message * Polishing the UI * optimization * rename variable * Update src/mlnet/AutoML/AutoMLEngine.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * new message * changed hhtp to https * added iteration num + 1 * change string name and add color to artifacts * change the message * build errors * added null checks * added exception messsages to log file * added exception messsages to log file * CLI ML.NET version upgrade (#345) * Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346) * CLI -- consume logs from AutoML SDK (#349) * Rename RunDetails --> RunDetail (#350) * command line api upgrade and progress bar rendering bug (#366) * added fix for all platforms progress bar * upgrade nuget * removed args from writeline * change in the version (#368) * fix few bugs in progressbar and verbosity (#374) * fix few bugs in progressbar and verbosity * removed unused name space * Fix for folders with space in it while generating project (#376) * support for folders with spaces * added support for paths with space * revert file * change name of var * remove spaces * SMAC fix for minimizing metrics (#363) * Formatting Regression metrics and progress bar display days. (#379) * added progress bar day display and fix regression metrics * fix formatting * added total time * formatted total time * change command name and add pbar message (#380) * change command name and add pbar message * fix tests * added aliases * duplicate alias * added another alias for task * UI missing features (#382) * added formatting changes * added accuracy specifically * downgrade the codepages (#384) * Change in project structure (#385) * initial changes * Change in project structure * correcting test * change variable name * fix tests * fix tests * fix more tests * fix codegen errors * adde log file message * changed name of args * change variable names * fix test * FileSizeBuckets in correct units (#387) * Minor telemetry change to log in correct units and make our life easier in the future * Use Ceiling instead of Round * changed order (#388) * prep work to transfer to ml.net (#389) * move test projects to top level test subdir * rename some projects to make naming consistent and make it build again * fix test project refs * Add AutoML components to build, fix issues related to that so it builds * fix test cases, remove AppInsights ref from AutoML (#3329) * [AutoML] disable netfx build leg for now (#3331) * disable netfx build leg for now * disable netfx build leg for now. * [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351) * [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364) * [AutoML] Minor changes to generated project in CLI based on feedback (#3371) * nitpicks for generated project * revert back the target framework * [AutoML] Migrate AutoML back to its own solution, w/ NuGet dependencies (#3373) * Migrate AutoML back to its own solution, w/ NuGet dependencies * build project updates; parameter name revert * dummy change * Revert "dummy change" This reverts commit 3e8574266f556a4d5b6805eb55b4d8b8b84cf355. * [AutoML] publish AutoML package (#3383) * publish AutoML package * Only leave automl and mlnet tests to run * publish AutoML package * Only leave automl and mlnet tests to run * fix build issues when ml.net is not building * bump version to 0.3 since that's the one we're going to ship for build (#3416) * [AutoML] temporarily disable all but x64 platforms -- don't want to do native builds and can't find a way around that with the current VSTS pipeline (#3420) * disable steps but keep phases to keep vsts build pipeline happy (#3423) * API docs for experimentation (#3484) * fixed path bug and regression metrics correction (#3504) * changed the casing of option alias as it conflicts with --help (#3554) * [AutoML] Generated project - FastTree nuget package inclusion dynamically (#3567) * added support for fast tree nuget pack inclusion in generated project * fix testcase * changed the tool name in telemetry message * dummy commit * remove space * dummy commit to trigger build * [AutoML] Add AutoML example code (#3458) * AutoML PipelineSuggester: don't recommend pipelines from first-stage trainers that failed (#3593) * InferColumns API: Validate all columns specified in column info exist in inferred data view (#3599) * [AutoML] AutoML SDK API: validate schema types of input IDataView (#3597) * [AutoML] If first three iterations all fail, short-circuit AutoML experiment (#3591) * mlnet CLI nupkg creation/signing (#3606) * mlnet CLI nupkg creation/signing * relmove includeinpackage from mlnet csproj * address PR comments -- some minor reshuffling of stuff * publish symbols for mlnet CLI * fix case in NLog.config * [AutoML] rename Auto to AutoML in namespace and nuget (#3609) * mlnet CLI nupkg creation/signing * [AutoML] take dependency on a specific ml.net version (#3610) * take dependency on a specific ml.net version * catch up to spelling fix for OptimizationTolerance * force a specific ml.net nuget version, fix typo (#3616) * [AutoML] Fix error handling in CLI. (#3618) * fix error handling * renaming variables * [AutoML] turn off line pragmas in .tt files to play nice with signing (#3617) * turn off line pragmas in .tt files to play nice with signing * dedupe tags * change the param name (#3619) * [AutoML] return null instead of null ref crash on Model property accessor (#3620) * return null instead of null ref crash on Model property accessor * [AutoML] Handling label column names which have space and exception logging (#3624) * fix case of label with space and exception logging * final handler * revert file * use Name instead of FullName for telemetry filename hash (#3633) * renamed classes (#3634) * change ML.NET dependency to 1.0 (#3639) [AutoML] undo pinning ML.NET dependency * set exploration time default in CLI to half hour (#3640) * [AutoML] step 2 of removing pinned nupkg versions (#3642) * InferColumns API that consumes label column index -- Only rename label column to 'Label' for headerless files (#3643) * [AutoML] Upgrade ml.net package in generated code (#3644) * upgrade the mlnet package in gen code * Update src/mlnet/Templates/Console/ModelProject.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/Templates/Console/ModelProject.tt Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * added spaces * [AutoML] Early stopping in CLI based on the exploration time (#3641) * early stopping in CLI * remove unused variables * change back to thread * remove sleep * fix review comments * remove ununsed usings * format message * collapse declaration * remove unused param * added environment.exit and removal of error message * correction in message * secs-> seconds * exit code * change value to 1 * reverse the declaration * [AutoML] Change wording for CouldNotFinshOnTime message (#3655) * set exploration time default in CLI to half hour * [AutoML] Change wording for CouldNotFinshOnTime message * [AutoML] Change wording for CouldNotFinshOnTime message * even better wording for CouldNotFinshOnTime * temp change to get around vsts publish failure (#3656) * [AutoML] bump version to 0.4.0 (#3658) * implement culture invariant strings (#3725) * reset culture (#3730) * [AutoML] Cross validation fixes; validate empty training / validation input data (#3794) * [AutoML] Enable style cop rules & resolve errors (#3823) * add task agnostic wrappers for autofit calls (#3860) * [AutoML] CLI telemetry rev (#3789) * delete automl .sln * CLI -- regenerate templated CS files (#3954) * [AutoML] Bump ML.NET package version to 1.2.0 in AutoML API and CLI; and AutoML package versions to 0.14.0 (#3958) * Build AutoML NuGet package (#3961) * Increment AutoML build version to 0.15.0 for preview. (#3968) * added culture independent parsing (#3731) * - convert tests to xunit - take project level dependency on ML.NET components instead of nuget - set up bestfriends relationship to ML.Core and remove some of the copies of util classes from AutoML.NET (more work needed to fully remove them, work item 4064) - misc build script changes to address PR comments * address issues only showing up in a couple configurations during CI build * fix cut&paste error * [AutoML] Bump version to ML.NET 1.3.1 in AutoML API and CLI and AutoML package version to 0.15.1 (#4071) * bumped version * change versions in nupkg * revert version bump in branch props * [AutoML] Fix for Exception thrown in cross val when one of the score equals infinity. (#4073) * bumped version * change versions in nupkg * revert version bump in branch props * added infinity fix * changes signing (#4079) * Addressed PR comments and build issues - sync block on creating test data file (failed intermittently) - removed classes we copied over from ML.Core and fixed their uses to de-dupe and use original ML.Core versions since we now have InternalsVisible and BestFriends - Fixed nupkg creation to use projects insted of public nuget version for AutoML - Fixed a bunch of unit tests that didn't actually test what they were supposed to test, while removing cut&past code and dependencies. - Few more misc small changes * minor nit - removed unused folder ref * Fix the .sln file for the right configurations. * Fix mistake in .sln file * test fixes and disable one test * fix tests, re-add AutoML samples csproj * bumped VS version to 16 in .sln, removed InternalsVisible for a dead assembly, removed unused references from AutoML test project * Updated docs to include PredictedLabel member (#4107) * Fixed build errors resulting from upgrade to VS2019 compilers * Added additional message describing the previous fix * Updated docs to include PredictedLabel member * Added CODEOWNERS file in the .github/ folder. (#4140) * Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository * Updated .github/CODEOWNERS with the team instead of individual reviewers * Added AutoML team reviewers (#4144) * Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository * Updated .github/CODEOWNERS with the team instead of individual reviewers * Added AutoML team reviwers to files owned by AutoML team * Added AutoML team reviwers to files owned by AutoML team * Removed two files that don't exist for AutoML team in CODEOWNERS * Build extension method to reload changes without specifying model name (#4146) * Image classification preview 2. (#4151) * Image classification preview 2. * PR feedback. * Add unit-test. * Add unit-test. * Add unit-test. * Add unit-test. * Use Path.Combine instead of Join. * fix test dataset path. * fix test dataset path. * Improve test. * Improve test. * Increase epochs in tests. * Disable test on Ubuntu. * Move test to its own project. * Move test to its own project. * Move test to its own project. * Move test to its own file. * cleanup. * Disable parallel execution of tensorflow tests. * PR feedback. * PR feedback. * PR feedback. * PR feedback. * Prevent TF test to execute in parallel. * PR feedback. * Build error. * clean up. * Added export functionality for LpNormNormalizingTransformer * Syncing upstream fork (#11) * Throw error on incorrect Label name in InferColumns API (#47) * Added sequential grouping of columns * reverted the file * addded infer columns label name checking * added column detection error * removed unsed usings * added quotes * replace Where with Any clause * replace Where with Any clause * Set Nullable Auto params to null values (#50) * Added sequential grouping of columns * reverted the file * added auto params as null * change to the update fields method * First public api propsal (#52) * Includes following 1) Final proposal for 0.1 public API surface 2) Prefeaturization 3) Splitting train data into train and validate when validation data is null 4) Providing end to end samples one each for regression, binaryclassification and multiclass classification * Incorporating code review feedbacks * Revert "Set Nullable Auto params to null values" (#53) * Revert "First public api propsal (#52)" This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7. * Revert "Set Nullable Auto params to null values (#50)" This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d. * AutoFit return type is now an IEnumerable (#55) AutoFit returns is now an IEnumerable - this enables many good things Implementing variety of early stopping criteria (See sample) Early discard of models that are no good. This improves memory usage efficiency. (See sample) No need to implement a callback to get results back Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample). Also templatized the return type for better type safety through out the code. * misc fixes & test additions, towards 0.1 release (#56) * Enable UnitTests on build server (#57) * 1) Making trainer name public (#62) 2) Fixing up samples to reflect it * Initial version of CLI tool for mlnet (#61) * added global tool initial project * removed unneccesary files, renamed files * refactoring and added base abstract classes for trainer generator * removed unused class * Added classes for transforms * added transform generate dummy classes * more refactoring, added first transform * more refactoring and added classes * changed the project structure * restructing added options class * sln changes * refactored options to different class: * added more logic for code generation of class * misc changes * reverted file * added commandline api package * reverted sample * added new command line api parser * added normalization of column names * Added command defaults and error message * implementation of all trainers * changed auto to null * added all transform generators * added error handling when args is empty and minor changes due to change in AutoML api names * changed the name of param * added new command line options and restructuring code * renamed proj file and added solution * Added code to generate usings, Fixed few bugs in the code * added validation to the command line options * changed project name * Bug fixes due to API change in AutoML * changed directory structure * added test framework and basic tests * added more tests * added improvements to template and error handling * renamed the estimator name * fixed test case * added comments * added headers * changed namespace and removed unneccesary properties from project * Revert "changed namespace and removed unneccesary properties from project" This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f. * fixed test cases and renamed namespaces * cleaned up proj file * added folder structure * added symbols/tokens for strings * added more tests * review comments * modified test cases * review comments * change in the exception message * normalized line endings * made method private static * simplified range building /optimization * minor fix * added header * added static methods in command where necessary * nit picks * made few methods static * review comments * nitpick * remove line pragmas * fix test case * Use better AutiFit overload and ignore Multiclass (#64) * Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65) * Added sequential grouping of columns * reverted the file * upgrade to v .10 and refactoring * added null check * fixed unit tests * review comments * removed the settings change * added regions * fixed unit tests * Upgrade ML.NET package to 0.10.0 (#70) * Change in template to accomodate new API of TextLoader (#72) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * Enable gated check for mlnet.tests (#79) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * added run-tests.proj and referred it in build.proj * CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83) * Added sequential grouping of columns * reverted the file * bug fixes, more logic to templates to support cross-validate * formatting and fix type in consolehelper * Added logic in templates * revert settings * benchmarking related changes (#63) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * fix fast forest learner (don't sweep over learning rate) (#88) * Made changes to Have non-calibrated scoring for binary classifiers (#86) * Added sequential grouping of columns * reverted the file * added calibration workaround * removed print probability * reverted settings * rev ColumnInference API: can take label index; rev output object types; add tests (#89) * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * publish nuget (#101) * use dotnet-internal-temp agent for internal build * use dotnet-internal feed * Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95) * Added sequential grouping of columns * reverted the file * fix usings for type convert * added transforms tests * review comments * When generating usings choose only distinct usings directives (#94) * Added sequential grouping of columns * reverted the file * Added code to have unique strings * refactoring * minor fix * minor fix * Autofit overloads + cancellation + progress callbacks 1) Introduce AutoFit overloads (basic and advanced) 2) AutoFit Cancellation 3) AutoFit progress callbacks * Default the kfolds to value 5 in CLI generated code (#115) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * remove file * added kfold param and defaulted to value * changed type * added for regression * Remove extra ; from generated code (#114) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * removed extra ; from generated code * removed file * fix unit tests * TimeoutInSeconds (#116) Specifying timeout in seconds instead of minutes * Added more command line args implementation to CLI tool and refactoring (#110) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * added git status * reverted change * added codegen options and refactoring * minor fixes' * renamed params, minor refactoring * added tests for commandline and refactoring * removed file * added back the test case * minor fixes * Update src/mlnet.Test/CommandLineTests.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * review comments * capitalize the first character * changed the name of test case * remove unused directives * Fail gracefully if unable to instantiate data view with swept parameters (#125) * gracefully fail if fail to parse a datai * rev * validate AutoFit 'Features' column must be of type R4 (#132) * Samples: exceptions / nits (#124) * Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121) * addded logging and helper methods * fixing code after merge * added resx files, added logger framework, added logging messages * added new options * added spacing * minor fixes * change command description * rename option, add headers, include new param in test * formatted * build fix * changed option name * Added NlogConfig file * added back config package * fix tests * added correct validation check (#137) * Use CreateTextLoader<T>(..) instead of CreateTextLoader(..) (#138) * added support to loaddata by class in the generated code * fix tests * changed CreateTextLoader to ReadFromTextFile method. (#140) * changed textloader to readfromtextfile method * formatting * exception fixes (#136) * infer purpose of hidden columns as 'ignore' (#142) * Added approval tests and bunch of refactoring of code and normalizing namespaces (#148) * changed textloader to readfromtextfile method * formatting * added approval tests and refactoring of code * removed few comments * API 2.0 skeleton (#149) Incorporating API review feedback * The CV code should come before the training when there is no test dataset in generated code (#151) * reorder cv code * build fix * fixed structure * Format the generated code + bunch of misc tasks (#152) * added formatting and minor changes for reordering cv * fixing the template * minor changes * formatting changes * fixed approval test * removed unused nuget * added missing value replacing * added test for new transform * fix test * Update src/mlnet/Templates/Console/MLCodeGen.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Sanitize the column names in CLI (#162) * added sanitization layer in CLI * fix test * changed exception.StackTrace to exception.ToString() * fix package name (#168) * Rev public API (#163) * Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153) * Fix minor version for the repository + remove Nlog config package (#171) * changed the minor version * removed the nlog config package * Added new test to columninfo and fixing up API (#178) * Make optimizing metric customizable and add trainer whitelist functionality (#172) * API rev (#181) * propagate root MLContext thru AutoML (instead of creating our own) (#182) * Enabling new command line args (#183) * fix package name * initial commit * added more commandline args * fixed tests * added headers * fix tests * fix test * rename 'AutoFitter' to 'Experiment' (#169) * added tests (#187) * rev InferColumns to accept ColumnInfo input param (#186) * Implement argument --has-header and change usage of dataset (#194) * added has header and fixed dataset and train dataset * fix tests * removed dummy command (#195) * Fix bug for regression and sanitize input label from user (#198) * removed dummy command * sanitize label and fix template * fix tests * Do not generate code concatenating columns when the dataset has a single feature column (#191) * Include some missed logging in the generated code. (#199) * added logging messages for generated code * added log messages * deleted file * cleaning up proj files (#185) * removed platform target * removed platform target * Some spaces and extra lines + bug in output path (#204) * nit picks * nit picks * fix test * accept label from user input and provide in generated code (#205) * Rev handling of weight / label columns (#203) * migrate to private ML.NET nuget for latest bug fixes (#131) * fix multiclass with nonstandard label (#207) * Multiclass nondefault label test (#208) * printing escaped chars + bug (#212) * delete unused internal samples (#211) * fix SMAC bug that causes multiclass sample to infinite loop (#209) * Rev user input validation for new API (#210) * added console message for exit and nit picks (#215) * exit when exception encountered (#216) * Seal API classes (and make EnableCaching internal) (#217) * Suggested sample nits (feel free to ask for any of these to be reverted) (#219) * User input column type validation (#218) * upgrade commandline and renaming (#221) * upgrade commandline and renaming * renaming fields * Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225) * CLI argument descriptions updated (#224) * CLI argument descriptions updated * No version in .csproj * added flag to disable training code (#227) * Exit if perfect model produced (#220) * removed header (#228) * removed header * added auto generated header * removed console read key (#229) * Fix model path in generated file (#230) * removed console read key * fix model path * fix test * reorder samples (#231) * remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233) * Null reference exception fix for finding best model when some runs have failed (#239) * samples fixes (#238) * fix for defaulting Averaged Perceptron # of iterations to 10 (#237) * Bug bash feedback Feb 27. API changes and sample changes (#240) * Bug bash feedback Feb 27. API changes Sample changes Exception fix * Samples / API rev from 2/27 bug bash feedback (#242) * changed the directory structure for generated project (#243) * changed the directory structure for generated project * changed test * upgraded commandline package * Fix test file locations on OSX (#235) * fix test file locations on OSX * changing to Path.Combine() * Additional Path.Combine() * Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt * Additional Path.Combine() * add back in double comparison fix * remove metrics agent NaN returns * test fix * test format fix * mock out path Thanks to @daholste for additional fixes! * upgrade to latest ML.NET public surface (#246) * Upgrade to ML.NET 0.11 (#247) * initial changes * fix lightgbm * changed normalize method * added tests * fix tests * fix test * Private preview final API changes (#250) * .NET framework design guidelines applied to public surface * WhitelistedTrainers -> Trainers * Add estimator to public API iteration result (#248) * LightGBM pipeline serialization fix (#251) * Change order that we search for TextLoader's parameters (#256) * CLI IFileInfo null exception fix (#254) * Averaged Perceptron pipeline serialization fix (#257) * Upgrade command-line-api and default folder name change (#258) * change in defautl folderName * upgrade command line * Update src/mlnet/Program.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * eliminate IFileInfo from CLI (#260) * Rev samples towards private preview; ignored columns fix (#259) * remove unused methods in consolehelper and nit picks in generated code (#261) * nit picks * change in console helper * fix tests * add space * fix tests * added nuget sources in generated csproj (#262) * added nuget sources in csproj * changed the structure in generated code * space * upgrade to mlnet 0.11 (#263) * Formatting CLI metrics (#264) Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits. * Add implementation of non -ova multi class trainers code gen (#267) * added non ova multi class learners * added tests * test cases * Add caching (#249) * AdvancedExperimentSettings sample nits (#265) * Add sampling key column (#268) * Initial work for multi-class classification support for CLI (#226) * Initial work for multi-class classification support for CLI * String updates * more strings * Whitelist non-OVA multi-class learners * Refactor the orchestration of AutoML calls (#272) * Do not auto-group columns with suggested purpose = 'Ignore' (#273) * Fix: during type inferencing, parse whitespace strings as NaN (#271) * Printing additional metrics in CLI for binary classification (#274) * Printing additional metrics in CLI for binary classification * Update src/mlnet/Utilities/ConsolePrinter.cs * Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269) * Print failed iterations in CLI (#275) * change the type to float from double (#277) * cache arg implementation in CLI (#280) * cache implementation * corrected the null case * added tests for all cases * Remove duplicate value-to-key mapping transform for multiclass string labels (#283) * Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286) * Implement ignore columns command line arg (#290) * normalize line endings * added --ignore-columns * null checks * unit tests * Print winning iteration and runtime in CLI (#288) * Print best metric and runtime * Print best metric and runtime * Line endings in AutoMLEngine.cs * Rename time column to duration to match Python SDK * Revert to MicroAccuracy and MacroAccuracy spellings * Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts * Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts * missed some files * Fix merge conflict * Update AutoMLEngine.cs * Add MacOS & Linux to CI; MacOS & Linux test fixes (#293) * MicroAccuracy as default for multi-class (#295) Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy. * Null exception for ignorecolumns in CLI (#294) * Null exception for ignorecolumns in CLI * Check if ignore-columns array has values (as the default is now a empty array) * Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296) * removed sln (#297) * Caching enabling in code gen part -2 (#298) * add * added caching codegen * support comma separated values for --ignore-columns (#300) * default initialization for ignore columns (#302) * default initialization * adde null check * Codegen for multiclass non-ova (#303) * changes to template * multicalss codegen * test cases * fix test cases * Generated Project new structure. (#305) * added new templates * writing files to disck * change path * added new templates * misisng braces * fix bugs * format code * added util methods for solution file creation and addition of projects to it * added extra packages to project files * new tests * added correct path for sln * build fix * fix build * include using system in prediction class (#307) * added using * fix test * Random number generator is not thread safe (#310) * Random number generator is not thread safe * Another local random generator * Missed a few references * Referncing AutoMlUtils.random instead of a local RNG * More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669 * Missed Random.cs * Fix multiclass code gen (#314) * compile error in codegen * removes scores printing * fix bugs * fix test * Fix compile error in codegen project (#319) * removed redundant code * fix test case * Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317) * Ova Multi class codegen support (#321) * dummy * multiova implementation * fix tests * remove inclusion list * fix tests and console helper * Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322) * Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination * test fixes * Console helper bug in generated code for multiclass (#323) * fix * fix test * looping perlogclass * fix test * Initial version of Progress bar impl and CLI UI experience (#325) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * Setting model directory to temp directory (#327) * Suggested changes to progress bar (#335) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * Rev Samples (#334) * Telemetry2 (#333) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * CLI telemetry implementation * Telemetry implementation * delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value * add headers, remove comments * one more header missing * Fix progress bar in linux/osx (#336) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * change from task to thread * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Mem leak fix (#328) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * there is still investigation to be done but this fix works and solves memory leak problems * minor refactor * Upgrade ML.NET package (#343) * Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287) * restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344) * Polishing the CLI UI part-1 (#338) * formatting of pbar message * Polishing the UI * optimization * rename variable * Update src/mlnet/AutoML/AutoMLEngine.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * new message * changed hhtp to https * added iteration num + 1 * change string name and add color to artifacts * change the message * build errors * added null checks * added exception messsages to log file * added exception messsages to log file * CLI ML.NET version upgrade (#345) * Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346) * CLI -- consume logs from AutoML SDK (#349) * Rename RunDetails --> RunDetail (#350) * command line api upgrade and progress bar rendering bug (#366) * added fix for all platforms progress bar * upgrade nuget * removed args from writeline * change in the version (#368) * fix few bugs in progressbar and verbosity (#374) * fix few bugs in progressbar and verbosity * removed unused name space * Fix for folders with space in it while generating project (#376) * support for folders with spaces * added support for paths with space * revert file * change name of var * remove spaces * SMAC fix for minimizing metrics (#363) * Formatting Regression metrics and progress bar display days. (#379) * added progress bar day display and fix regression metrics * fix formatting * added total time * formatted total time * change command name and add pbar message (#380) * change command name and add pbar message * fix tests * added aliases * duplicate alias * added another alias for task * UI missing features (#382) * added formatting changes * added accuracy specifically * downgrade the codepages (#384) * Change in project structure (#385) * initial changes * Change in project structure * correcting test * change variable name * fix tests * fix tests * fix more tests * fix codegen errors * adde log file message * changed name of args * change variable names * fix test * FileSizeBuckets in correct units (#387) * Minor telemetry change to log in correct units and make our life easier in the future * Use Ceiling instead of Round * changed order (#388) * prep work to transfer to ml.net (#389) * move test projects to top level test subdir * rename some projects to make naming consistent and make it build again * fix test project refs * Add AutoML components to build, fix issues related to that so it builds * fix test cases, remove AppInsights ref from AutoML (#3329) * [AutoML] disable netfx build leg for now (#3331) * disable netfx build leg for now * disable netfx build leg for now. * [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351) * [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364) * [AutoML] Minor changes to generated project in CLI based on feedback (#3371) * nitpicks for generated project * revert back the target framework * [AutoML] Migrate AutoML back to its own s…

* Fixed build errors resulting from upgrade to VS2019 compilers * Added additional message describing the previous fix * Syncing upstream fork (#10) * Throw error on incorrect Label name in InferColumns API (#47) * Added sequential grouping of columns * reverted the file * addded infer columns label name checking * added column detection error * removed unsed usings * added quotes * replace Where with Any clause * replace Where with Any clause * Set Nullable Auto params to null values (#50) * Added sequential grouping of columns * reverted the file * added auto params as null * change to the update fields method * First public api propsal (#52) * Includes following 1) Final proposal for 0.1 public API surface 2) Prefeaturization 3) Splitting train data into train and validate when validation data is null 4) Providing end to end samples one each for regression, binaryclassification and multiclass classification * Incorporating code review feedbacks * Revert "Set Nullable Auto params to null values" (#53) * Revert "First public api propsal (#52)" This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7. * Revert "Set Nullable Auto params to null values (#50)" This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d. * AutoFit return type is now an IEnumerable (#55) AutoFit returns is now an IEnumerable - this enables many good things Implementing variety of early stopping criteria (See sample) Early discard of models that are no good. This improves memory usage efficiency. (See sample) No need to implement a callback to get results back Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample). Also templatized the return type for better type safety through out the code. * misc fixes & test additions, towards 0.1 release (#56) * Enable UnitTests on build server (#57) * 1) Making trainer name public (#62) 2) Fixing up samples to reflect it * Initial version of CLI tool for mlnet (#61) * added global tool initial project * removed unneccesary files, renamed files * refactoring and added base abstract classes for trainer generator * removed unused class * Added classes for transforms * added transform generate dummy classes * more refactoring, added first transform * more refactoring and added classes * changed the project structure * restructing added options class * sln changes * refactored options to different class: * added more logic for code generation of class * misc changes * reverted file * added commandline api package * reverted sample * added new command line api parser * added normalization of column names * Added command defaults and error message * implementation of all trainers * changed auto to null * added all transform generators * added error handling when args is empty and minor changes due to change in AutoML api names * changed the name of param * added new command line options and restructuring code * renamed proj file and added solution * Added code to generate usings, Fixed few bugs in the code * added validation to the command line options * changed project name * Bug fixes due to API change in AutoML * changed directory structure * added test framework and basic tests * added more tests * added improvements to template and error handling * renamed the estimator name * fixed test case * added comments * added headers * changed namespace and removed unneccesary properties from project * Revert "changed namespace and removed unneccesary properties from project" This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f. * fixed test cases and renamed namespaces * cleaned up proj file * added folder structure * added symbols/tokens for strings * added more tests * review comments * modified test cases * review comments * change in the exception message * normalized line endings * made method private static * simplified range building /optimization * minor fix * added header * added static methods in command where necessary * nit picks * made few methods static * review comments * nitpick * remove line pragmas * fix test case * Use better AutiFit overload and ignore Multiclass (#64) * Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65) * Added sequential grouping of columns * reverted the file * upgrade to v .10 and refactoring * added null check * fixed unit tests * review comments * removed the settings change * added regions * fixed unit tests * Upgrade ML.NET package to 0.10.0 (#70) * Change in template to accomodate new API of TextLoader (#72) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * Enable gated check for mlnet.tests (#79) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * added run-tests.proj and referred it in build.proj * CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83) * Added sequential grouping of columns * reverted the file * bug fixes, more logic to templates to support cross-validate * formatting and fix type in consolehelper * Added logic in templates * revert settings * benchmarking related changes (#63) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * fix fast forest learner (don't sweep over learning rate) (#88) * Made changes to Have non-calibrated scoring for binary classifiers (#86) * Added sequential grouping of columns * reverted the file * added calibration workaround * removed print probability * reverted settings * rev ColumnInference API: can take label index; rev output object types; add tests (#89) * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * publish nuget (#101) * use dotnet-internal-temp agent for internal build * use dotnet-internal feed * Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95) * Added sequential grouping of columns * reverted the file * fix usings for type convert * added transforms tests * review comments * When generating usings choose only distinct usings directives (#94) * Added sequential grouping of columns * reverted the file * Added code to have unique strings * refactoring * minor fix * minor fix * Autofit overloads + cancellation + progress callbacks 1) Introduce AutoFit overloads (basic and advanced) 2) AutoFit Cancellation 3) AutoFit progress callbacks * Default the kfolds to value 5 in CLI generated code (#115) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * remove file * added kfold param and defaulted to value * changed type * added for regression * Remove extra ; from generated code (#114) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * removed extra ; from generated code * removed file * fix unit tests * TimeoutInSeconds (#116) Specifying timeout in seconds instead of minutes * Added more command line args implementation to CLI tool and refactoring (#110) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * added git status * reverted change * added codegen options and refactoring * minor fixes' * renamed params, minor refactoring * added tests for commandline and refactoring * removed file * added back the test case * minor fixes * Update src/mlnet.Test/CommandLineTests.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * review comments * capitalize the first character * changed the name of test case * remove unused directives * Fail gracefully if unable to instantiate data view with swept parameters (#125) * gracefully fail if fail to parse a datai * rev * validate AutoFit 'Features' column must be of type R4 (#132) * Samples: exceptions / nits (#124) * Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121) * addded logging and helper methods * fixing code after merge * added resx files, added logger framework, added logging messages * added new options * added spacing * minor fixes * change command description * rename option, add headers, include new param in test * formatted * build fix * changed option name * Added NlogConfig file * added back config package * fix tests * added correct validation check (#137) * Use CreateTextLoader<T>(..) instead of CreateTextLoader(..) (#138) * added support to loaddata by class in the generated code * fix tests * changed CreateTextLoader to ReadFromTextFile method. (#140) * changed textloader to readfromtextfile method * formatting * exception fixes (#136) * infer purpose of hidden columns as 'ignore' (#142) * Added approval tests and bunch of refactoring of code and normalizing namespaces (#148) * changed textloader to readfromtextfile method * formatting * added approval tests and refactoring of code * removed few comments * API 2.0 skeleton (#149) Incorporating API review feedback * The CV code should come before the training when there is no test dataset in generated code (#151) * reorder cv code * build fix * fixed structure * Format the generated code + bunch of misc tasks (#152) * added formatting and minor changes for reordering cv * fixing the template * minor changes * formatting changes * fixed approval test * removed unused nuget * added missing value replacing * added test for new transform * fix test * Update src/mlnet/Templates/Console/MLCodeGen.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Sanitize the column names in CLI (#162) * added sanitization layer in CLI * fix test * changed exception.StackTrace to exception.ToString() * fix package name (#168) * Rev public API (#163) * Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153) * Fix minor version for the repository + remove Nlog config package (#171) * changed the minor version * removed the nlog config package * Added new test to columninfo and fixing up API (#178) * Make optimizing metric customizable and add trainer whitelist functionality (#172) * API rev (#181) * propagate root MLContext thru AutoML (instead of creating our own) (#182) * Enabling new command line args (#183) * fix package name * initial commit * added more commandline args * fixed tests * added headers * fix tests * fix test * rename 'AutoFitter' to 'Experiment' (#169) * added tests (#187) * rev InferColumns to accept ColumnInfo input param (#186) * Implement argument --has-header and change usage of dataset (#194) * added has header and fixed dataset and train dataset * fix tests * removed dummy command (#195) * Fix bug for regression and sanitize input label from user (#198) * removed dummy command * sanitize label and fix template * fix tests * Do not generate code concatenating columns when the dataset has a single feature column (#191) * Include some missed logging in the generated code. (#199) * added logging messages for generated code * added log messages * deleted file * cleaning up proj files (#185) * removed platform target * removed platform target * Some spaces and extra lines + bug in output path (#204) * nit picks * nit picks * fix test * accept label from user input and provide in generated code (#205) * Rev handling of weight / label columns (#203) * migrate to private ML.NET nuget for latest bug fixes (#131) * fix multiclass with nonstandard label (#207) * Multiclass nondefault label test (#208) * printing escaped chars + bug (#212) * delete unused internal samples (#211) * fix SMAC bug that causes multiclass sample to infinite loop (#209) * Rev user input validation for new API (#210) * added console message for exit and nit picks (#215) * exit when exception encountered (#216) * Seal API classes (and make EnableCaching internal) (#217) * Suggested sample nits (feel free to ask for any of these to be reverted) (#219) * User input column type validation (#218) * upgrade commandline and renaming (#221) * upgrade commandline and renaming * renaming fields * Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225) * CLI argument descriptions updated (#224) * CLI argument descriptions updated * No version in .csproj * added flag to disable training code (#227) * Exit if perfect model produced (#220) * removed header (#228) * removed header * added auto generated header * removed console read key (#229) * Fix model path in generated file (#230) * removed console read key * fix model path * fix test * reorder samples (#231) * remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233) * Null reference exception fix for finding best model when some runs have failed (#239) * samples fixes (#238) * fix for defaulting Averaged Perceptron # of iterations to 10 (#237) * Bug bash feedback Feb 27. API changes and sample changes (#240) * Bug bash feedback Feb 27. API changes Sample changes Exception fix * Samples / API rev from 2/27 bug bash feedback (#242) * changed the directory structure for generated project (#243) * changed the directory structure for generated project * changed test * upgraded commandline package * Fix test file locations on OSX (#235) * fix test file locations on OSX * changing to Path.Combine() * Additional Path.Combine() * Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt * Additional Path.Combine() * add back in double comparison fix * remove metrics agent NaN returns * test fix * test format fix * mock out path Thanks to @daholste for additional fixes! * upgrade to latest ML.NET public surface (#246) * Upgrade to ML.NET 0.11 (#247) * initial changes * fix lightgbm * changed normalize method * added tests * fix tests * fix test * Private preview final API changes (#250) * .NET framework design guidelines applied to public surface * WhitelistedTrainers -> Trainers * Add estimator to public API iteration result (#248) * LightGBM pipeline serialization fix (#251) * Change order that we search for TextLoader's parameters (#256) * CLI IFileInfo null exception fix (#254) * Averaged Perceptron pipeline serialization fix (#257) * Upgrade command-line-api and default folder name change (#258) * change in defautl folderName * upgrade command line * Update src/mlnet/Program.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * eliminate IFileInfo from CLI (#260) * Rev samples towards private preview; ignored columns fix (#259) * remove unused methods in consolehelper and nit picks in generated code (#261) * nit picks * change in console helper * fix tests * add space * fix tests * added nuget sources in generated csproj (#262) * added nuget sources in csproj * changed the structure in generated code * space * upgrade to mlnet 0.11 (#263) * Formatting CLI metrics (#264) Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits. * Add implementation of non -ova multi class trainers code gen (#267) * added non ova multi class learners * added tests * test cases * Add caching (#249) * AdvancedExperimentSettings sample nits (#265) * Add sampling key column (#268) * Initial work for multi-class classification support for CLI (#226) * Initial work for multi-class classification support for CLI * String updates * more strings * Whitelist non-OVA multi-class learners * Refactor the orchestration of AutoML calls (#272) * Do not auto-group columns with suggested purpose = 'Ignore' (#273) * Fix: during type inferencing, parse whitespace strings as NaN (#271) * Printing additional metrics in CLI for binary classification (#274) * Printing additional metrics in CLI for binary classification * Update src/mlnet/Utilities/ConsolePrinter.cs * Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269) * Print failed iterations in CLI (#275) * change the type to float from double (#277) * cache arg implementation in CLI (#280) * cache implementation * corrected the null case * added tests for all cases * Remove duplicate value-to-key mapping transform for multiclass string labels (#283) * Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286) * Implement ignore columns command line arg (#290) * normalize line endings * added --ignore-columns * null checks * unit tests * Print winning iteration and runtime in CLI (#288) * Print best metric and runtime * Print best metric and runtime * Line endings in AutoMLEngine.cs * Rename time column to duration to match Python SDK * Revert to MicroAccuracy and MacroAccuracy spellings * Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts * Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts * missed some files * Fix merge conflict * Update AutoMLEngine.cs * Add MacOS & Linux to CI; MacOS & Linux test fixes (#293) * MicroAccuracy as default for multi-class (#295) Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy. * Null exception for ignorecolumns in CLI (#294) * Null exception for ignorecolumns in CLI * Check if ignore-columns array has values (as the default is now a empty array) * Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296) * removed sln (#297) * Caching enabling in code gen part -2 (#298) * add * added caching codegen * support comma separated values for --ignore-columns (#300) * default initialization for ignore columns (#302) * default initialization * adde null check * Codegen for multiclass non-ova (#303) * changes to template * multicalss codegen * test cases * fix test cases * Generated Project new structure. (#305) * added new templates * writing files to disck * change path * added new templates * misisng braces * fix bugs * format code * added util methods for solution file creation and addition of projects to it * added extra packages to project files * new tests * added correct path for sln * build fix * fix build * include using system in prediction class (#307) * added using * fix test * Random number generator is not thread safe (#310) * Random number generator is not thread safe * Another local random generator * Missed a few references * Referncing AutoMlUtils.random instead of a local RNG * More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669 * Missed Random.cs * Fix multiclass code gen (#314) * compile error in codegen * removes scores printing * fix bugs * fix test * Fix compile error in codegen project (#319) * removed redundant code * fix test case * Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317) * Ova Multi class codegen support (#321) * dummy * multiova implementation * fix tests * remove inclusion list * fix tests and console helper * Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322) * Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination * test fixes * Console helper bug in generated code for multiclass (#323) * fix * fix test * looping perlogclass * fix test * Initial version of Progress bar impl and CLI UI experience (#325) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * Setting model directory to temp directory (#327) * Suggested changes to progress bar (#335) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * Rev Samples (#334) * Telemetry2 (#333) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * CLI telemetry implementation * Telemetry implementation * delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value * add headers, remove comments * one more header missing * Fix progress bar in linux/osx (#336) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * change from task to thread * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Mem leak fix (#328) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * there is still investigation to be done but this fix works and solves memory leak problems * minor refactor * Upgrade ML.NET package (#343) * Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287) * restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344) * Polishing the CLI UI part-1 (#338) * formatting of pbar message * Polishing the UI * optimization * rename variable * Update src/mlnet/AutoML/AutoMLEngine.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * new message * changed hhtp to https * added iteration num + 1 * change string name and add color to artifacts * change the message * build errors * added null checks * added exception messsages to log file * added exception messsages to log file * CLI ML.NET version upgrade (#345) * Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346) * CLI -- consume logs from AutoML SDK (#349) * Rename RunDetails --> RunDetail (#350) * command line api upgrade and progress bar rendering bug (#366) * added fix for all platforms progress bar * upgrade nuget * removed args from writeline * change in the version (#368) * fix few bugs in progressbar and verbosity (#374) * fix few bugs in progressbar and verbosity * removed unused name space * Fix for folders with space in it while generating project (#376) * support for folders with spaces * added support for paths with space * revert file * change name of var * remove spaces * SMAC fix for minimizing metrics (#363) * Formatting Regression metrics and progress bar display days. (#379) * added progress bar day display and fix regression metrics * fix formatting * added total time * formatted total time * change command name and add pbar message (#380) * change command name and add pbar message * fix tests * added aliases * duplicate alias * added another alias for task * UI missing features (#382) * added formatting changes * added accuracy specifically * downgrade the codepages (#384) * Change in project structure (#385) * initial changes * Change in project structure * correcting test * change variable name * fix tests * fix tests * fix more tests * fix codegen errors * adde log file message * changed name of args * change variable names * fix test * FileSizeBuckets in correct units (#387) * Minor telemetry change to log in correct units and make our life easier in the future * Use Ceiling instead of Round * changed order (#388) * prep work to transfer to ml.net (#389) * move test projects to top level test subdir * rename some projects to make naming consistent and make it build again * fix test project refs * Add AutoML components to build, fix issues related to that so it builds * fix test cases, remove AppInsights ref from AutoML (#3329) * [AutoML] disable netfx build leg for now (#3331) * disable netfx build leg for now * disable netfx build leg for now. * [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351) * [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364) * [AutoML] Minor changes to generated project in CLI based on feedback (#3371) * nitpicks for generated project * revert back the target framework * [AutoML] Migrate AutoML back to its own solution, w/ NuGet dependencies (#3373) * Migrate AutoML back to its own solution, w/ NuGet dependencies * build project updates; parameter name revert * dummy change * Revert "dummy change" This reverts commit 3e8574266f556a4d5b6805eb55b4d8b8b84cf355. * [AutoML] publish AutoML package (#3383) * publish AutoML package * Only leave automl and mlnet tests to run * publish AutoML package * Only leave automl and mlnet tests to run * fix build issues when ml.net is not building * bump version to 0.3 since that's the one we're going to ship for build (#3416) * [AutoML] temporarily disable all but x64 platforms -- don't want to do native builds and can't find a way around that with the current VSTS pipeline (#3420) * disable steps but keep phases to keep vsts build pipeline happy (#3423) * API docs for experimentation (#3484) * fixed path bug and regression metrics correction (#3504) * changed the casing of option alias as it conflicts with --help (#3554) * [AutoML] Generated project - FastTree nuget package inclusion dynamically (#3567) * added support for fast tree nuget pack inclusion in generated project * fix testcase * changed the tool name in telemetry message * dummy commit * remove space * dummy commit to trigger build * [AutoML] Add AutoML example code (#3458) * AutoML PipelineSuggester: don't recommend pipelines from first-stage trainers that failed (#3593) * InferColumns API: Validate all columns specified in column info exist in inferred data view (#3599) * [AutoML] AutoML SDK API: validate schema types of input IDataView (#3597) * [AutoML] If first three iterations all fail, short-circuit AutoML experiment (#3591) * mlnet CLI nupkg creation/signing (#3606) * mlnet CLI nupkg creation/signing * relmove includeinpackage from mlnet csproj * address PR comments -- some minor reshuffling of stuff * publish symbols for mlnet CLI * fix case in NLog.config * [AutoML] rename Auto to AutoML in namespace and nuget (#3609) * mlnet CLI nupkg creation/signing * [AutoML] take dependency on a specific ml.net version (#3610) * take dependency on a specific ml.net version * catch up to spelling fix for OptimizationTolerance * force a specific ml.net nuget version, fix typo (#3616) * [AutoML] Fix error handling in CLI. (#3618) * fix error handling * renaming variables * [AutoML] turn off line pragmas in .tt files to play nice with signing (#3617) * turn off line pragmas in .tt files to play nice with signing * dedupe tags * change the param name (#3619) * [AutoML] return null instead of null ref crash on Model property accessor (#3620) * return null instead of null ref crash on Model property accessor * [AutoML] Handling label column names which have space and exception logging (#3624) * fix case of label with space and exception logging * final handler * revert file * use Name instead of FullName for telemetry filename hash (#3633) * renamed classes (#3634) * change ML.NET dependency to 1.0 (#3639) [AutoML] undo pinning ML.NET dependency * set exploration time default in CLI to half hour (#3640) * [AutoML] step 2 of removing pinned nupkg versions (#3642) * InferColumns API that consumes label column index -- Only rename label column to 'Label' for headerless files (#3643) * [AutoML] Upgrade ml.net package in generated code (#3644) * upgrade the mlnet package in gen code * Update src/mlnet/Templates/Console/ModelProject.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/Templates/Console/ModelProject.tt Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * added spaces * [AutoML] Early stopping in CLI based on the exploration time (#3641) * early stopping in CLI * remove unused variables * change back to thread * remove sleep * fix review comments * remove ununsed usings * format message * collapse declaration * remove unused param * added environment.exit and removal of error message * correction in message * secs-> seconds * exit code * change value to 1 * reverse the declaration * [AutoML] Change wording for CouldNotFinshOnTime message (#3655) * set exploration time default in CLI to half hour * [AutoML] Change wording for CouldNotFinshOnTime message * [AutoML] Change wording for CouldNotFinshOnTime message * even better wording for CouldNotFinshOnTime * temp change to get around vsts publish failure (#3656) * [AutoML] bump version to 0.4.0 (#3658) * implement culture invariant strings (#3725) * reset culture (#3730) * [AutoML] Cross validation fixes; validate empty training / validation input data (#3794) * [AutoML] Enable style cop rules & resolve errors (#3823) * add task agnostic wrappers for autofit calls (#3860) * [AutoML] CLI telemetry rev (#3789) * delete automl .sln * CLI -- regenerate templated CS files (#3954) * [AutoML] Bump ML.NET package version to 1.2.0 in AutoML API and CLI; and AutoML package versions to 0.14.0 (#3958) * Build AutoML NuGet package (#3961) * Increment AutoML build version to 0.15.0 for preview. (#3968) * added culture independent parsing (#3731) * - convert tests to xunit - take project level dependency on ML.NET components instead of nuget - set up bestfriends relationship to ML.Core and remove some of the copies of util classes from AutoML.NET (more work needed to fully remove them, work item 4064) - misc build script changes to address PR comments * address issues only showing up in a couple configurations during CI build * fix cut&paste error * [AutoML] Bump version to ML.NET 1.3.1 in AutoML API and CLI and AutoML package version to 0.15.1 (#4071) * bumped version * change versions in nupkg * revert version bump in branch props * [AutoML] Fix for Exception thrown in cross val when one of the score equals infinity. (#4073) * bumped version * change versions in nupkg * revert version bump in branch props * added infinity fix * changes signing (#4079) * Addressed PR comments and build issues - sync block on creating test data file (failed intermittently) - removed classes we copied over from ML.Core and fixed their uses to de-dupe and use original ML.Core versions since we now have InternalsVisible and BestFriends - Fixed nupkg creation to use projects insted of public nuget version for AutoML - Fixed a bunch of unit tests that didn't actually test what they were supposed to test, while removing cut&past code and dependencies. - Few more misc small changes * minor nit - removed unused folder ref * Fix the .sln file for the right configurations. * Fix mistake in .sln file * test fixes and disable one test * fix tests, re-add AutoML samples csproj * bumped VS version to 16 in .sln, removed InternalsVisible for a dead assembly, removed unused references from AutoML test project * Updated docs to include PredictedLabel member (#4107) * Fixed build errors resulting from upgrade to VS2019 compilers * Added additional message describing the previous fix * Updated docs to include PredictedLabel member * Added CODEOWNERS file in the .github/ folder. (#4140) * Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository * Updated .github/CODEOWNERS with the team instead of individual reviewers * Added AutoML team reviewers (#4144) * Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository * Updated .github/CODEOWNERS with the team instead of individual reviewers * Added AutoML team reviwers to files owned by AutoML team * Added AutoML team reviwers to files owned by AutoML team * Removed two files that don't exist for AutoML team in CODEOWNERS * Build extension method to reload changes without specifying model name (#4146) * Image classification preview 2. (#4151) * Image classification preview 2. * PR feedback. * Add unit-test. * Add unit-test. * Add unit-test. * Add unit-test. * Use Path.Combine instead of Join. * fix test dataset path. * fix test dataset path. * Improve test. * Improve test. * Increase epochs in tests. * Disable test on Ubuntu. * Move test to its own project. * Move test to its own project. * Move test to its own project. * Move test to its own file. * cleanup. * Disable parallel execution of tensorflow tests. * PR feedback. * PR feedback. * PR feedback. * PR feedback. * Prevent TF test to execute in parallel. * PR feedback. * Build error. * clean up. * Syncing upstream fork (#11) * Throw error on incorrect Label name in InferColumns API (#47) * Added sequential grouping of columns * reverted the file * addded infer columns label name checking * added column detection error * removed unsed usings * added quotes * replace Where with Any clause * replace Where with Any clause * Set Nullable Auto params to null values (#50) * Added sequential grouping of columns * reverted the file * added auto params as null * change to the update fields method * First public api propsal (#52) * Includes following 1) Final proposal for 0.1 public API surface 2) Prefeaturization 3) Splitting train data into train and validate when validation data is null 4) Providing end to end samples one each for regression, binaryclassification and multiclass classification * Incorporating code review feedbacks * Revert "Set Nullable Auto params to null values" (#53) * Revert "First public api propsal (#52)" This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7. * Revert "Set Nullable Auto params to null values (#50)" This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d. * AutoFit return type is now an IEnumerable (#55) AutoFit returns is now an IEnumerable - this enables many good things Implementing variety of early stopping criteria (See sample) Early discard of models that are no good. This improves memory usage efficiency. (See sample) No need to implement a callback to get results back Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample). Also templatized the return type for better type safety through out the code. * misc fixes & test additions, towards 0.1 release (#56) * Enable UnitTests on build server (#57) * 1) Making trainer name public (#62) 2) Fixing up samples to reflect it * Initial version of CLI tool for mlnet (#61) * added global tool initial project * removed unneccesary files, renamed files * refactoring and added base abstract classes for trainer generator * removed unused class * Added classes for transforms * added transform generate dummy classes * more refactoring, added first transform * more refactoring and added classes * changed the project structure * restructing added options class * sln changes * refactored options to different class: * added more logic for code generation of class * misc changes * reverted file * added commandline api package * reverted sample * added new command line api parser * added normalization of column names * Added command defaults and error message * implementation of all trainers * changed auto to null * added all transform generators * added error handling when args is empty and minor changes due to change in AutoML api names * changed the name of param * added new command line options and restructuring code * renamed proj file and added solution * Added code to generate usings, Fixed few bugs in the code * added validation to the command line options * changed project name * Bug fixes due to API change in AutoML * changed directory structure * added test framework and basic tests * added more tests * added improvements to template and error handling * renamed the estimator name * fixed test case * added comments * added headers * changed namespace and removed unneccesary properties from project * Revert "changed namespace and removed unneccesary properties from project" This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f. * fixed test cases and renamed namespaces * cleaned up proj file * added folder structure * added symbols/tokens for strings * added more tests * review comments * modified test cases * review comments * change in the exception message * normalized line endings * made method private static * simplified range building /optimization * minor fix * added header * added static methods in command where necessary * nit picks * made few methods static * review comments * nitpick * remove line pragmas * fix test case * Use better AutiFit overload and ignore Multiclass (#64) * Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65) * Added sequential grouping of columns * reverted the file * upgrade to v .10 and refactoring * added null check * fixed unit tests * review comments * removed the settings change * added regions * fixed unit tests * Upgrade ML.NET package to 0.10.0 (#70) * Change in template to accomodate new API of TextLoader (#72) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * Enable gated check for mlnet.tests (#79) * Added sequential grouping of columns * reverted the file * changed to new API of Text Loader * changed signature * added params for taking additional settings * changes to codegen params * refactoring of templates and fixing errors * added run-tests.proj and referred it in build.proj * CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83) * Added sequential grouping of columns * reverted the file * bug fixes, more logic to templates to support cross-validate * formatting and fix type in consolehelper * Added logic in templates * revert settings * benchmarking related changes (#63) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * fix fast forest learner (don't sweep over learning rate) (#88) * Made changes to Have non-calibrated scoring for binary classifiers (#86) * Added sequential grouping of columns * reverted the file * added calibration workaround * removed print probability * reverted settings * rev ColumnInference API: can take label index; rev output object types; add tests (#89) * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * publish nuget (#101) * use dotnet-internal-temp agent for internal build * use dotnet-internal feed * Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95) * Added sequential grouping of columns * reverted the file * fix usings for type convert * added transforms tests * review comments * When generating usings choose only distinct usings directives (#94) * Added sequential grouping of columns * reverted the file * Added code to have unique strings * refactoring * minor fix * minor fix * Autofit overloads + cancellation + progress callbacks 1) Introduce AutoFit overloads (basic and advanced) 2) AutoFit Cancellation 3) AutoFit progress callbacks * Default the kfolds to value 5 in CLI generated code (#115) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * remove file * added kfold param and defaulted to value * changed type * added for regression * Remove extra ; from generated code (#114) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * removed extra ; from generated code * removed file * fix unit tests * TimeoutInSeconds (#116) Specifying timeout in seconds instead of minutes * Added more command line args implementation to CLI tool and refactoring (#110) * Added sequential grouping of columns * reverted the file * Set up CI with Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * Update azure-pipelines.yml for Azure Pipelines * added git status * reverted change * added codegen options and refactoring * minor fixes' * renamed params, minor refactoring * added tests for commandline and refactoring * removed file * added back the test case * minor fixes * Update src/mlnet.Test/CommandLineTests.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * review comments * capitalize the first character * changed the name of test case * remove unused directives * Fail gracefully if unable to instantiate data view with swept parameters (#125) * gracefully fail if fail to parse a datai * rev * validate AutoFit 'Features' column must be of type R4 (#132) * Samples: exceptions / nits (#124) * Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121) * addded logging and helper methods * fixing code after merge * added resx files, added logger framework, added logging messages * added new options * added spacing * minor fixes * change command description * rename option, add headers, include new param in test * formatted * build fix * changed option name * Added NlogConfig file * added back config package * fix tests * added correct validation check (#137) * Use CreateTextLoader<T>(..) instead of CreateTextLoader(..) (#138) * added support to loaddata by class in the generated code * fix tests * changed CreateTextLoader to ReadFromTextFile method. (#140) * changed textloader to readfromtextfile method * formatting * exception fixes (#136) * infer purpose of hidden columns as 'ignore' (#142) * Added approval tests and bunch of refactoring of code and normalizing namespaces (#148) * changed textloader to readfromtextfile method * formatting * added approval tests and refactoring of code * removed few comments * API 2.0 skeleton (#149) Incorporating API review feedback * The CV code should come before the training when there is no test dataset in generated code (#151) * reorder cv code * build fix * fixed structure * Format the generated code + bunch of misc tasks (#152) * added formatting and minor changes for reordering cv * fixing the template * minor changes * formatting changes * fixed approval test * removed unused nuget * added missing value replacing * added test for new transform * fix test * Update src/mlnet/Templates/Console/MLCodeGen.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Sanitize the column names in CLI (#162) * added sanitization layer in CLI * fix test * changed exception.StackTrace to exception.ToString() * fix package name (#168) * Rev public API (#163) * Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153) * Fix minor version for the repository + remove Nlog config package (#171) * changed the minor version * removed the nlog config package * Added new test to columninfo and fixing up API (#178) * Make optimizing metric customizable and add trainer whitelist functionality (#172) * API rev (#181) * propagate root MLContext thru AutoML (instead of creating our own) (#182) * Enabling new command line args (#183) * fix package name * initial commit * added more commandline args * fixed tests * added headers * fix tests * fix test * rename 'AutoFitter' to 'Experiment' (#169) * added tests (#187) * rev InferColumns to accept ColumnInfo input param (#186) * Implement argument --has-header and change usage of dataset (#194) * added has header and fixed dataset and train dataset * fix tests * removed dummy command (#195) * Fix bug for regression and sanitize input label from user (#198) * removed dummy command * sanitize label and fix template * fix tests * Do not generate code concatenating columns when the dataset has a single feature column (#191) * Include some missed logging in the generated code. (#199) * added logging messages for generated code * added log messages * deleted file * cleaning up proj files (#185) * removed platform target * removed platform target * Some spaces and extra lines + bug in output path (#204) * nit picks * nit picks * fix test * accept label from user input and provide in generated code (#205) * Rev handling of weight / label columns (#203) * migrate to private ML.NET nuget for latest bug fixes (#131) * fix multiclass with nonstandard label (#207) * Multiclass nondefault label test (#208) * printing escaped chars + bug (#212) * delete unused internal samples (#211) * fix SMAC bug that causes multiclass sample to infinite loop (#209) * Rev user input validation for new API (#210) * added console message for exit and nit picks (#215) * exit when exception encountered (#216) * Seal API classes (and make EnableCaching internal) (#217) * Suggested sample nits (feel free to ask for any of these to be reverted) (#219) * User input column type validation (#218) * upgrade commandline and renaming (#221) * upgrade commandline and renaming * renaming fields * Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225) * CLI argument descriptions updated (#224) * CLI argument descriptions updated * No version in .csproj * added flag to disable training code (#227) * Exit if perfect model produced (#220) * removed header (#228) * removed header * added auto generated header * removed console read key (#229) * Fix model path in generated file (#230) * removed console read key * fix model path * fix test * reorder samples (#231) * remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233) * Null reference exception fix for finding best model when some runs have failed (#239) * samples fixes (#238) * fix for defaulting Averaged Perceptron # of iterations to 10 (#237) * Bug bash feedback Feb 27. API changes and sample changes (#240) * Bug bash feedback Feb 27. API changes Sample changes Exception fix * Samples / API rev from 2/27 bug bash feedback (#242) * changed the directory structure for generated project (#243) * changed the directory structure for generated project * changed test * upgraded commandline package * Fix test file locations on OSX (#235) * fix test file locations on OSX * changing to Path.Combine() * Additional Path.Combine() * Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt * Additional Path.Combine() * add back in double comparison fix * remove metrics agent NaN returns * test fix * test format fix * mock out path Thanks to @daholste for additional fixes! * upgrade to latest ML.NET public surface (#246) * Upgrade to ML.NET 0.11 (#247) * initial changes * fix lightgbm * changed normalize method * added tests * fix tests * fix test * Private preview final API changes (#250) * .NET framework design guidelines applied to public surface * WhitelistedTrainers -> Trainers * Add estimator to public API iteration result (#248) * LightGBM pipeline serialization fix (#251) * Change order that we search for TextLoader's parameters (#256) * CLI IFileInfo null exception fix (#254) * Averaged Perceptron pipeline serialization fix (#257) * Upgrade command-line-api and default folder name change (#258) * change in defautl folderName * upgrade command line * Update src/mlnet/Program.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * eliminate IFileInfo from CLI (#260) * Rev samples towards private preview; ignored columns fix (#259) * remove unused methods in consolehelper and nit picks in generated code (#261) * nit picks * change in console helper * fix tests * add space * fix tests * added nuget sources in generated csproj (#262) * added nuget sources in csproj * changed the structure in generated code * space * upgrade to mlnet 0.11 (#263) * Formatting CLI metrics (#264) Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits. * Add implementation of non -ova multi class trainers code gen (#267) * added non ova multi class learners * added tests * test cases * Add caching (#249) * AdvancedExperimentSettings sample nits (#265) * Add sampling key column (#268) * Initial work for multi-class classification support for CLI (#226) * Initial work for multi-class classification support for CLI * String updates * more strings * Whitelist non-OVA multi-class learners * Refactor the orchestration of AutoML calls (#272) * Do not auto-group columns with suggested purpose = 'Ignore' (#273) * Fix: during type inferencing, parse whitespace strings as NaN (#271) * Printing additional metrics in CLI for binary classification (#274) * Printing additional metrics in CLI for binary classification * Update src/mlnet/Utilities/ConsolePrinter.cs * Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269) * Print failed iterations in CLI (#275) * change the type to float from double (#277) * cache arg implementation in CLI (#280) * cache implementation * corrected the null case * added tests for all cases * Remove duplicate value-to-key mapping transform for multiclass string labels (#283) * Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286) * Implement ignore columns command line arg (#290) * normalize line endings * added --ignore-columns * null checks * unit tests * Print winning iteration and runtime in CLI (#288) * Print best metric and runtime * Print best metric and runtime * Line endings in AutoMLEngine.cs * Rename time column to duration to match Python SDK * Revert to MicroAccuracy and MacroAccuracy spellings * Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts * Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts * missed some files * Fix merge conflict * Update AutoMLEngine.cs * Add MacOS & Linux to CI; MacOS & Linux test fixes (#293) * MicroAccuracy as default for multi-class (#295) Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy. * Null exception for ignorecolumns in CLI (#294) * Null exception for ignorecolumns in CLI * Check if ignore-columns array has values (as the default is now a empty array) * Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296) * removed sln (#297) * Caching enabling in code gen part -2 (#298) * add * added caching codegen * support comma separated values for --ignore-columns (#300) * default initialization for ignore columns (#302) * default initialization * adde null check * Codegen for multiclass non-ova (#303) * changes to template * multicalss codegen * test cases * fix test cases * Generated Project new structure. (#305) * added new templates * writing files to disck * change path * added new templates * misisng braces * fix bugs * format code * added util methods for solution file creation and addition of projects to it * added extra packages to project files * new tests * added correct path for sln * build fix * fix build * include using system in prediction class (#307) * added using * fix test * Random number generator is not thread safe (#310) * Random number generator is not thread safe * Another local random generator * Missed a few references * Referncing AutoMlUtils.random instead of a local RNG * More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669 * Missed Random.cs * Fix multiclass code gen (#314) * compile error in codegen * removes scores printing * fix bugs * fix test * Fix compile error in codegen project (#319) * removed redundant code * fix test case * Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317) * Ova Multi class codegen support (#321) * dummy * multiova implementation * fix tests * remove inclusion list * fix tests and console helper * Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322) * Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination * test fixes * Console helper bug in generated code for multiclass (#323) * fix * fix test * looping perlogclass * fix test * Initial version of Progress bar impl and CLI UI experience (#325) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * Setting model directory to temp directory (#327) * Suggested changes to progress bar (#335) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * Rev Samples (#334) * Telemetry2 (#333) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * CLI telemetry implementation * Telemetry implementation * delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value * add headers, remove comments * one more header missing * Fix progress bar in linux/osx (#336) * progressbar * added progressbar and refactoring * reverted * revert sign assembly * added headers and removed exception rethrow * bug fixes and updates to UI * added friendly name printing for metric * formatting * change from task to thread * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Mem leak fix (#328) * Create test.txt * Create test.txt * changes needed for benchmarking * forgot one file * merge conflict fix * fix build break * back out my version of the fix for Label column issue and fix the original fix * bogus file removal * undo SuggestedPipeline change * remove labelCol from pipeline suggester * fix build break * rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline) * tweak queue in vsts-ci.yml * there is still investigation to be done but this fix works and solves memory leak problems * minor refactor * Upgrade ML.NET package (#343) * Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287) * restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344) * Polishing the CLI UI part-1 (#338) * formatting of pbar message * Polishing the UI * optimization * rename variable * Update src/mlnet/AutoML/AutoMLEngine.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com> * new message * changed hhtp to https * added iteration num + 1 * change string name and add color to artifacts * change the message * build errors * added null checks * added exception messsages to log file * added exception messsages to log file * CLI ML.NET version upgrade (#345) * Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346) * CLI -- consume logs from AutoML SDK (#349) * Rename RunDetails --> RunDetail (#350) * command line api upgrade and progress bar rendering bug (#366) * added fix for all platforms progress bar * upgrade nuget * removed args from writeline * change in the version (#368) * fix few bugs in progressbar and verbosity (#374) * fix few bugs in progressbar and verbosity * removed unused name space * Fix for folders with space in it while generating project (#376) * support for folders with spaces * added support for paths with space * revert file * change name of var * remove spaces * SMAC fix for minimizing metrics (#363) * Formatting Regression metrics and progress bar display days. (#379) * added progress bar day display and fix regression metrics * fix formatting * added total time * formatted total time * change command name and add pbar message (#380) * change command name and add pbar message * fix tests * added aliases * duplicate alias * added another alias for task * UI missing features (#382) * added formatting changes * added accuracy specifically * downgrade the codepages (#384) * Change in project structure (#385) * initial changes * Change in project structure * correcting test * change variable name * fix tests * fix tests * fix more tests * fix codegen errors * adde log file message * changed name of args * change variable names * fix test * FileSizeBuckets in correct units (#387) * Minor telemetry change to log in correct units and make our life easier in the future * Use Ceiling instead of Round * changed order (#388) * prep work to transfer to ml.net (#389) * move test projects to top level test subdir * rename some projects to make naming consistent and make it build again * fix test project refs * Add AutoML components to build, fix issues related to that so it builds * fix test cases, remove AppInsights ref from AutoML (#3329) * [AutoML] disable netfx build leg for now (#3331) * disable netfx build leg for now * disable netfx build leg for now. * [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351) * [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364) * [AutoML] Minor changes to generated project in CLI based on feedback (#3371) * nitpicks for generated project * revert back the target framework * [AutoML] Migrate AutoML back to its own solution, w/ NuGet dependencies (#3373) * Migrate AutoML back to its own solut…

Added GitHubLabeler and GettingStarted

2ff3a88

- Iris sample - Sentiment Analysis sample - TaxiFare sample -GitHub issues classification sample

OliaG requested review from codemzs, eerhardt, terrajobst, asthana86 and shauheen May 22, 2018 05:52

asthana86 approved these changes May 22, 2018

View reviewed changes

asthana86 requested a review from danmoseley May 22, 2018 06:34

justinormont reviewed May 22, 2018

View reviewed changes

shauheen closed this May 22, 2018

shauheen reopened this May 22, 2018

markusweimer reviewed May 22, 2018

View reviewed changes

TomFinley reviewed May 22, 2018

View reviewed changes

justinormont reviewed May 22, 2018

View reviewed changes

eerhardt reviewed May 22, 2018

View reviewed changes

TomFinley suggested changes May 23, 2018

View reviewed changes

TomFinley reviewed May 23, 2018

View reviewed changes

mandyshieh and others added 2 commits May 23, 2018 10:30

Added Block Size Checks for ParquetLoader (#120)

7a5b303

* added ceiling to mathutils and have parquetloader default to sequential reading instead of throwing * factor out sequence creation and add check for overflow in MathUtils

terrajobst suggested changes May 23, 2018

View reviewed changes

yaeldMS and others added 3 commits May 23, 2018 11:30

CV macro with stratification column doesn't work (#213)

73d894b

* Reduce number of hash bits in stratification column and add a unit test. * Address PR comments.

Address review comments

c90ad8a

Fixes build errors caused by spaces in the project path (#196)

76393f4

eerhardt mentioned this pull request May 23, 2018

Taxi fare dataset is almost 50MB #206

Closed

Ivanidzo4ka and others added 8 commits May 23, 2018 16:22

switch housing dataset to wine (#170)

d51321c

* replace housing uci dataset to wine quality

Reorganized datasets

8658ce0

Moved datasets from tests to examples Removed model files

Added GitHubLabeler and GettingStarted

6f0b329

- Iris sample - Sentiment Analysis sample - TaxiFare sample -GitHub issues classification sample

Address review comments

8ac5db8

Reorganized datasets

5ca7cd5

Moved datasets from tests to examples Removed model files

Merge branch 'samples' of https://github.com/dotnet/machinelearning i…

9afcccc

…nto samples

Rebase with master, remove merge conflicts

7cbd7d7

Fixed unit tests

8509d48

Add README for GettingStarted sln

a66ed98

TomFinley reviewed May 24, 2018

View reviewed changes

OliaG closed this May 24, 2018

OliaG mentioned this pull request May 24, 2018

Update samples #238

Closed

justinormont mentioned this pull request Sep 11, 2018

Hot linking to a UCI dataset #889

Closed

Dmitry-A pushed a commit to Dmitry-A/machinelearning that referenced this pull request Apr 12, 2019

Fix bug for regression and sanitize input label from user (dotnet#198)

b3f980b

* removed dummy command * sanitize label and fix template * fix tests

Dmitry-A pushed a commit to Dmitry-A/machinelearning that referenced this pull request Aug 22, 2019

Fix bug for regression and sanitize input label from user (dotnet#198)

a89bc48

* removed dummy command * sanitize label and fix template * fix tests

ghost locked as resolved and limited conversation to collaborators Mar 30, 2022

		@@ -0,0 +1,2 @@
		<wpf:ResourceDictionary xml:space="preserve" xmlns:x="http://schemas.microsoft.com/winfx/2006/xaml" xmlns:s="clr-namespace:System;assembly=mscorlib" xmlns:ss="urn:shemas-jetbrains-com:settings-storage-xaml" xmlns:wpf="http://schemas.microsoft.com/winfx/2006/xaml/presentation">
		<s:String x:Key="/Default/CodeInspection/CSharpLanguageProject/LanguageLevel/@EntryValue">CSharp72</s:String></wpf:ResourceDictionary>

		@@ -0,0 +1,11 @@
		MICROSOFT PROVIDES THE DATASETS ON AN "AS IS" BASIS. MICROSOFT MAKES NO WARRANTIES, EXPRESS OR IMPLIED, GUARANTEES OR CONDITIONS WITH RESPECT TO YOUR USE OF THE DATASETS. TO THE EXTENT PERMITTED UNDER YOUR LOCAL LAW, MICROSOFT DISCLAIMS ALL LIABILITY FOR ANY DAMAGES OR LOSSES, INLCUDING DIRECT, CONSEQUENTIAL, SPECIAL, INDIRECT, INCIDENTAL OR PUNITIVE, RESULTING FROM YOUR USE OF THE DATASETS.

		@@ -0,0 +1,1000 @@
		A very, very, very slow-moving, aimless movie about a distressed, drifting young man. 0
		Not sure who was more lost - the flat characters or the audience, nearly half of whom walked out. 0

		@@ -0,0 +1,191 @@
		# IDV File Format

		This document describes ML.NET's Binary dataview file format, version 1.1.1.5

Added samples: GitHubLabeler and GettingStarted #198

Added samples: GitHubLabeler and GettingStarted #198

Conversation

OliaG commented May 22, 2018

asthana86 left a comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

justinormont May 22, 2018 • edited Loading

Choose a reason for hiding this comment

shauheen commented May 22, 2018

markusweimer left a comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

TomFinley May 22, 2018 • edited Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

OliaG May 23, 2018 • edited by terrajobst Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

TomFinley commented May 22, 2018

TomFinley left a comment • edited Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

KrzysztofCwalina commented May 24, 2018

OliaG commented May 24, 2018

Choose a reason for hiding this comment

justinormont May 22, 2018 •

edited

Loading

TomFinley May 22, 2018 •

edited

Loading

OliaG May 23, 2018 •

edited by terrajobst

Loading

TomFinley left a comment •

edited

Loading