Introduction to Machine learning

IntroductiontoMachinelearningIntroductiontomachinelearningbyQuentindeLaroussilhe-@Underflow404MachinelearningAmachinelearningalgorithmisanalgorithmlearningtoaccomplishataskbyobservingdata.●Usedoncomplextaskswhereit’shardtodevelopalgorithmswithhandcrafted-rules●ExploitspatternsinobserveddataandextractrulesautomaticallyFieldsofapplication●Computervision●Speechrecognition●Financialanalysis●Searchengines●Ads-targeting●Contentsuggestion●Self-drivingcars●Assistants●etc...Example:objectdetectionBigvariationinvisualfeatures:●Shape●Background●Size/positionClassifyinganobjectinapictureisnotaneasytask.Example:objectdetection●Learnfromannotatedcorpusofexamples(adataset)toclassifyunknownimagesamongdifferentobjecttypes●Observeimagestolearnpatterns●Lotofdataavailable(i.e:ImageNetdataset)●Verygooderrorrates(CNN)GeneralconceptsTypesofMLalgorithmsSupervisedLearnafunctionbyobservingexamplescontainingtheinputandtheexpectedoutput.●Classification●RegressionUnsupervisedFindunderliningrelationsindatabyobservingtherawdataonly(withouttheexpectedoutput).●Clustering●DimensionalityreductionTrainingsetClassificationvsRegressionRegressionLearnafunctionmappinganinputelementtoarealvalue.i.e:PredictthetemperatureoftomorrowgivensomemeteosignalsClassificationLearnafunctionmappinganinputelementtoaclass(withinafinitesetofpossibleclasses).i.e:Predicttheweatheroftomorrow:{sunny,cloudy,rainy}givensomemeteosignalsRegressionClassificationClusteringAclusteringalgorithmseparatedifferentobserveddatapointsinsimilargroups(clusters).Wedonotknowthelabelsduringtraining.Cluster1Cluster3Cluster2ReinforcementlearningLearntheoptimalbehaviorforanagentinanenvironmenttomaximizeagivengoal.Examples:●Driveacaronaroadandminimizethecollisionrisk●Playvideo-games●ChoosethepositionofadsonawebsitetomaximizethenumberofclicksFeatureextractionThefirststepinamachinelearningprocessistoextractusefulvaluesfromthedata(calledfeatures).Thegoalistoextracttheinformationusefulforthetaskwewanttolearn.Examples:●Stockmarkettime-serie→[openingprice,closingprice,lowest,highest]●Image→Imagewithedgesfiltered●Document→bag-of-wordModelisationprocessknearestneighborsk-nearestneighbors●Classificationandregressionmodel●Supervisedlearning:wehaveannotatedexamples●Weclassifyanewexamplebasedonthelabelsofhis“nearestneighbors”●kisthenumberofneighborstakeninconsiderationk-nearestneighborsToclassifyapoint:Welookthek-nearestneighbors(herek=5)andwedoamajorityvote.Thispointhas3redneighborsand2blueneighbors,itwillbeclassifiedasred.k-nearestneighbors●Ndatapoints●Requireadistancefunctionbetweenpoints●Regression(averagethevalueofthek-nearestneighbors)●Classification(majorityvoteofthek-nearestneighbors)k-nearestneighbors:effectofk●kisthenumberofneighborstakeninconsideration●Ifk=1○Theaccuracyonthetrainingsetis100%○Itmightnotgeneralizeonnewdata●Ifk1○Theaccuracyonthetrainingsetmightnotbe100%○Itmightgeneralizebetteronunseendatak-nearestneighbors:weightedversionInthecaseofunbalancedrepartitionbetweenclasseswecangiveweightstotheexamples.●Theweightofaunderrepresentedclasswillbesethigh.●Theweightofaoverrepresentedclasswillbesetlow.Whenwedothemajorityvote,wetaketheweightinconsideration:●Forclassificationwedoaweightedvote.●Forregressionwedoaweightedaverage.Decisiontrees,randomforestsDecisiontreeDecisiontree●Decisiontreespartitionthefeaturespacebysplittingthedata●LearningthedecisiontreeconsistsinfindingtheorderandthesplitcriterionforeachnodeDecisiontree●Thedecisiontreelearningisparametrizedbythemethodforchoosingthesplitsandthemaximumheight●Ifthemaximumheightisbigenough,alltheexamplesofthetrainingdatawouldbecorrectlyclassified:overfitting.Decisiontree:entropymetric●S:Thedatasetsbeforethesplit●X:Setofexistingclasses●p(x):ProportionofelementsinclassxtothenumberofelementsinS●A:Thesplitcriterion●T:ThedatasetscreatedbythesplitEntropy:Ateachstepwecreatethenodebysplittingwiththecriterionwiththehighestinformationgain.Randomforests●Whenthedepthofadecisiontreeisgrowingtheerroronvalidationdatatendstoincreasealot:highvariance●OnewaytoexploitalotofdataistotrainmultipledecisiontreesandaveragethemAlgorithm:●SelectNpointsinthetrainingdataandkfeatures(usuallysqrt(p))●Learnanewdecisiontree●StopwhenwehaveenoughtreesClustering-kmeansClusteringwithk-means●Clusteringalgorithm●Requireadistancefunctionbetweenpoints●kisthenumberofclusterthealgorithmwillfindClusteringwithk-meansObjective:Dividethedatasetinksetsbyminimizingthewithin-clustersumofsquares(sumofdistancesofeachpointoftheclustertothecenterofthecluster)WhereSarethesetswearelearningandμthemeanoftheseti.iClusteringwithk-meansGradientdescentGradientdescent1.DefineamodeldependingonW(theparametersofthemodel)2.Definealossfunctionthatquantifytheerrorthemodeldoesonthetrainingdata:○convergence○themodelisgoodenough○yourspentallyourmoneyLinearregressionLinearregression:IntroductionInsupervisedlearning,wehaveexamplesofl

Introduction to Machine learning

免费阅读已结束，点击付费阅读剩下 ... 页

阅读已结束，您可以下载文档离线阅读

2020年银行年度考核表个人工作总结（共4篇）

2020年银行专题组织生活会情况报告

2020年银行筑牢支付安全防线喜迎幸福美好新年主题宣传活动工作报告总结

2020年银行支行经营管理合规性自查工作报告总结

2020年银行支行意识形态工作总结

2020年银行营销金融知识进万家活动工作总结

2020年银行意识形态工作总结

2020年银行信用社扫黑除恶专项斗争开展汇报工作总结

2020年银行信托业务部业务培训总结

2020年银行网点转型操作风险自评估报告

相关文档

相关搜索