聚类分析外文文献

11ClusterAnalysisThenexttwochaptersaddressclassiﬁcationissuesfromtwovaryingperspectives.Whenconsideringgroupsofobjectsinamultivariatedataset,twosituationscanarise.Givenadatasetcontainingmeasurementsonindividuals,insomecaseswewanttoseeifsomenaturalgroupsorclassesofindividualsexist,andinothercases,wewanttoclassifytheindividualsaccordingtoasetofexistinggroups.Clusteranalysisdevelopstoolsandmethodsconcerningtheformercase,thatis,givenadatamatrixcontainingmultivariatemeasurementsonalargenumberofindividuals(orobjects),theobjectiveistobuildsomenaturalsubgroupsorclustersofindividuals.Thisisdonebygroupingindividualsthatare“similar”accordingtosomeappropriatecriterion.Oncetheclustersareobtained,itisgenerallyusefultodescribeeachgroupusingsomedescriptivetoolfromChapters1,8or9tocreateabetterunderstandingofthediﬀerencesthatexistamongtheformulatedgroups.Clusteranalysisisappliedinmanyﬁeldssuchasthenaturalsciences,themedicalsciences,economics,marketing,etc.Inmarketing,forinstance,itisusefultobuildanddescribethediﬀerentsegmentsofamarketfromasurveyonpotentialconsumers.Aninsurancecompany,ontheotherhand,mightbeinterestedinthedistinctionamongclassesofpotentialcustomerssothatitcanderiveoptimalpricesforitsservices.Otherexamplesareprovidedbelow.DiscriminantanalysispresentedinChapter12addressestheotherissueofclassiﬁcation.Itfocusesonsituationswherethediﬀerentgroupsareknownapriori.Decisionrulesareprovidedinclassifyingamultivariateobservationintooneoftheknowngroups.Section11.1statestheproblemofclusteranalysiswherethecriterionchosentomeasurethesimilarityamongobjectsclearlyplaysanimportantrole.Section11.2showshowtopreciselymeasuretheproximitybetweenobjects.Finally,Section11.3providessomealgorithms.Wewillconcentrateonhierarchicalalgorithmsonlywherethenumberofclustersisnotknowninadvance.11.1TheProblemClusteranalysisisasetoftoolsforbuildinggroups(clusters)frommultivariatedataobjects.Theaimistoconstructgroupswithhomogeneouspropertiesoutofheterogeneouslargesamples.Thegroupsorclustersshouldbeashomogeneousaspossibleandthediﬀerencesamongthevariousgroupsaslargeaspossible.Clusteranalysiscanbedividedintotwofundamentalsteps.1.Choiceofaproximitymeasure:Onecheckseachpairofobservations(objects)forthesimilarityoftheirvalues.Asimilarity(proximity)measureisdeﬁnedtomeasurethe“closeness”oftheobjects.The“closer”theyare,themorehomogeneoustheyare.2.Choiceofgroup-buildingalgorithm:Onthebasisoftheproximitymeasurestheobjectsassignedtogroupssothatdiﬀerencesbetweengroupsbecomelargeandobservationsinagroupbecomeascloseaspossible.30211ClusterAnalysisInmarketing,forexmaple,clusteranalysisisusedtoselecttestmarkets.Otherapplicationsincludetheclassiﬁcationofcompaniesaccordingtotheirorganizationalstructures,technolo-giesandtypes.Inpsychology,clusteranalysisisusedtoﬁndtypesofpersonalitiesonthebasisofquestionnaires.Inarchaeology,itisappliedtoclassifyartobjectsindiﬀerenttimeperiods.Otherscientiﬁcbranchesthatuseclusteranalysisaremedicine,sociology,linguis-ticsandbiology.Ineachcaseaheterogeneoussampleofobjectsareanalyzedwiththeaimtoidentifyhomogeneoussubgroups.11.2TheProximitybetweenObjectsThestartingpointofaclusteranalysisisadatamatrixX(n×p)withnmeasurements(objects)ofpvariables.Theproximity(similarity)amongobjectsisdescribedbyamatrixD(n×n)D=⎛⎜⎜⎜⎜⎜⎜⎜⎜⎜⎝d11d12.........d1n...d22.......................................dn1dn2.........dnn⎞⎟⎟⎟⎟⎟⎟⎟⎟⎟⎠.(11.1)ThematrixDcontainsmeasuresofsimilarityordissimilarityamongthenobjects.Ifthevaluesdijaredistances,thentheymeasuredissimilarity.Thegreaterthedistance,thelesssimilararetheobjects.Ifthevaluesdijareproximitymeasures,thentheoppositeistrue,i.e.,thegreatertheproximityvalue,themoresimilararetheobjects.Adistancematrix,forexample,couldbedeﬁnedbytheL2-norm:dij=xi−xj2,wherexiandxjdenotetherowsofthedatamatrixX.Distanceandsimilarityareofcoursedual.Ifdijisadistance,thendij=maxi,j{dij}−dijisaproximitymeasure.Thenatureoftheobservationsplaysanimportantroleinthechoiceofproximitymeasure.Nominalvalues(likebinaryvariables)leadingeneraltoproximityvalues,whereasmetricvalueslead(ingeneral)todistancematrices.WeﬁrstpresentpossibilitiesforDinthebinarycaseandthenconsiderthecontinuouscase.11.2TheProximitybetweenObjects303SimilarityofobjectswithbinarystructureInordertomeasurethesimilaritybetweenobjectswealwayscomparepairsofobservations(xi,xj)wherexi=(xi1,...,xip),xj=(xj1,...,xjp),andxik,xjk∈{0,1}.Obviouslytherearefourcases:xik=xjk=1,xik=0,xjk=1,xik=1,xjk=0,xik=xjk=0.NameδλDeﬁnitionJaccard01a1a1+a2+a3Tanimoto12a1+a4a1+2(a2+a3)+a4SimpleMatching(M)11a1+a4pRusselandRao(RR)––a1pDice00.52a12a1+(a2+a3)Kulczynski––a1a2+a3Table11.2.Thecommonsimilaritycoeﬃcients.Deﬁnea1=pk=1I(xik=xjk=1),a2=pk=1I(xik=0,xjk=1),a3=pk=1I(xik=1,xjk=0),a4=pk=1I(xik=xjk=0).30411ClusterAnalysisNotethateacha,=1,...,4,dependsonthepair(xi,xj).Thefollowingproximitymeasuresareusedinpractice:dij=a1+δa4a1+δa4+λ(a2+a3)(11.2)whereδandλareweightingfactors.Table11.2showssomesimilaritymeasuresforgivenweightingfactors.Thesemeasuresprovidealternativewaysofweightingmismatchingsandpositive(presenceofaco

聚类分析外文文献

免费阅读已结束，点击付费阅读剩下 ... 页

阅读已结束，您可以下载文档离线阅读

Pklebi经济管理类常用期刊目录

供应链模式下的OEM交货期策略优化问题研究

机械专业英语9687948394

建筑施工安全检查标准JGJ59-

(建筑篇)施工组织设计001

国际金融-外汇、汇率与汇率制度

第五章金融衍生工具-互换

我国甲醇发展现状及未来预测doc6)(1)

业务报文内容及格式管理办法(试行)

品牌和品牌价值

相关文档

相关搜索