Chapter 4 An Introduction to Hidden Markov Models

Chapter4AnIntroductiontoHiddenMarkovModelsforBiologicalSequencesbyAndersKroghCenterforBiologicalSequenceAnalysisTechnicalUniversityofDenmarkBuilding206,2800Lyngby,DenmarkPhone:+4545252471Fax:+4545934808E-mail:krogh@cbs.dtu.dkInComputationalMethodsinMolecularBiology,editedbyS.L.Salzberg,D.B.SearlsandS.Kasif,pages45-63.Elsevier,1998.1Contents4AnIntroductiontoHiddenMarkovModelsforBiologicalSequences14.1Introduction.............................34.2FromregularexpressionstoHMMs................44.3ProﬁleHMMs............................84.3.1Pseudocounts........................104.3.2Searchingadatabase....................114.3.3Modelestimation......................124.4HMMsforgeneﬁnding.......................134.4.1Signalsensors.......................144.4.2Codingregions.......................154.4.3Combiningthemodels...................164.5Furtherreading...........................1924.1IntroductionVeryefﬁcientprogramsforsearchingatextforacombinationofwordsareavail-ableonmanycomputers.Thesamemethodscanbeusedforsearchingforpatternsinbiologicalsequences,butoftentheyfail.Thisisbecausebiological‘spelling’ismuchmoresloppythanEnglishspelling:proteinswiththesamefunctionfromtwodifferentorganismsarealmostcertainlyspelleddifferently,thatis,thetwoaminoacidsequencesdiffer.Itisnotrarethattwosuchhomologoussequenceshavelessthan30%identicalaminoacids.SimilarlyinDNAmanyinterestingsig-nalsvarygreatlyevenwithinthesamegenome.Somewell-knownexamplesareribosomebindingsitesandsplicesites,butthelistislong.Fortunatelythereareusuallystillsomesubtlesimilaritiesbetweentwosuchsequences,andtheques-tionishowtodetectthesesimilarities.Thevariationinafamilyofsequencescanbedescribedstatistically,andthisisthebasisformostmethodsusedinbiologicalsequenceanalysis,see[1]forapresentationofsomeofthesestatisticalapproaches.Forpairwisealignments,forinstance,theprobabilitythatacertainresiduemutatestoanotherresidueisusedinasubstitutionmatrix,suchasoneofthePAMmatrices.ForﬁndingpatternsinDNA,e.g.splicesites,somesortofweightmatrixisveryoftenused,whichissimplyapositionspeciﬁcscorecalculatedfromthefrequenciesofthefournucleotidesatallthepositionsinsomeknownexamples.Similarly,methodsforﬁndinggenesuse,almostwithoutexception,thestatisticsofcodonsordicodonsinsomeformorother.AhiddenMarkovmodel(HMM)isastatisticalmodel,whichisverywellsuitedformanytasksinmolecularbiology,althoughtheyhavebeenmostlyde-velopedforspeechrecognitionsincetheearly1970s,see[2]forhistoricaldetails.ThemostpopularuseoftheHMMinmolecularbiologyisasa‘probabilisticpro-ﬁle’ofaproteinfamily,whichiscalledaproﬁleHMM.Fromafamilyofproteins(orDNA)aproﬁleHMMcanbemadeforsearchingadatabaseforothermem-bersofthefamily.TheseproﬁleHMMsresembletheproﬁle[3]andweightmatrixmethods[4,5],andprobablythemaincontributionisthattheproﬁleHMMtreatsgapsinasystematicway.TheHMMcanbeappliedtoothertypesofproblems.Itisparticularlywellsuitedforproblemswithasimple‘grammaticalstructure,’suchasgeneﬁnding.Ingeneﬁndingseveralsignalsmustberecognizedandcombinedintoapredictionofexonsandintrons,andthepredictionmustconformtovariousrulestomakeitareasonablegeneprediction.AnHMMcancombinerecognitionofthesignals,anditcanbemadesuchthatthepredictionsalwaysfollowtherulesofagene.SincemuchoftheliteratureonHMMsisalittlehardtoreadformanybiol-ogists,Iwillattemptinthischaptertogiveanon-mathematicalintroductiontoHMMs.Whereasthelittlebiologicalbackgroundneededistakenforgranted,I3havetriedtoexplainHMMsatalevelthatalmostanyonecanfollow.FirstHMMsareintroducedbyanexampleandthenproﬁleHMMsaredescribed.ThenanHMMforﬁndingeukaryoticgenesissketched,andﬁnallypointerstothelitera-turearegiven.4.2FromregularexpressionstoHMMsMostreadershavenodoubtcomeacrossregularexpressionsatsomepoint,andmanyprobablyusethemquitealot.Regularexpressionsareusedinmanypro-grams,inparticularonUnixcomputers.Inprogramslikeawk,grep,sed,andperl,regularexpressionscanbeusedforsearchingtextﬁlesforapattern.Withgrepforinstance,youcansearchaﬁleforalllinescontaining‘C.elegans’or‘Caenorhab-ditiselegans’withtheregularexpression‘’.ThiswillmatchanylinecontainingaCfollowedbyanynumberoflower-caselettersor‘.’,thenaspaceandthenelegans.Regularexpressionscanalsobeusedtocharacterizeproteinfamilies,whichisthebasisforthePROSITEdatabase[6].Usingregularexpressionsisaveryelegantandefﬁcientwaytosearchforsomeproteinfamilies,butdifﬁcultforother.Asalreadymentionedinthein-troduction,thedifﬁcultiesarisebecauseproteinspellingismuchmorefreethanEnglishspelling.Thereforetheregularexpressionssometimesneedtobeverybroadandcomplex.ImagineaDNAmotiflikethis:!$#!$!!$##!$#!$(IuseDNAonlybecauseofthesmallernumberoflettersthanforaminoacids).Aregularexpressionforthisis[AT][CG][AC][ACGT]*A[TG][GC],meaningthattheﬁrstpositionisAorT,thesecondCorG,andsoforth.Theterm‘[ACGT]*’meansthatanyofthefourletterscanoccuranynumberoftimes.Theproblemwiththeaboveregularexpressionisthatitdoesnotinanywaydistinguishbetweenthehighlyimplausiblesequence!#%!##whichhastheexception

Chapter 4 An Introduction to Hidden Markov Models

免费阅读已结束，点击付费阅读剩下 ... 页

阅读已结束，您可以下载文档离线阅读

开源ERP优势分析

国投昔阳综合信息化平台培训文档

PCB电镀沉铜药水控制工艺参数

万科广场设计分享

第十章建筑工程进度管理

中国人寿保险股份有限公司投资连结保险投资账户 XXXX 年半年信息_

关于西安交通大学第十一届教学成果奖推荐及评审工作的通知

OTC零售终端及商务管理(doc15)(1)

药品价格策略

HDX8000终端操作使用

相关文档

相关搜索

Chapter 4 An Introduction to Hidden Markov Models

免费阅读已结束，点击付费阅读剩下 ... 页

阅读已结束，您可以下载文档离线阅读

开源ERP优势分析

国投昔阳综合信息化平台培训文档

PCB电镀沉铜药水控制工艺参数

万科广场设计分享

第十章建筑工程进度管理

中国人寿保险股份有限公司 投资连结保险投资账户 XXXX 年半年信息_

关于西安交通大学第十一届教学成果奖推荐及评审工作的通知

OTC零售终端及商务管理(doc15)(1)

药品价格策略

HDX8000终端操作使用

相关文档

相关搜索

中国人寿保险股份有限公司投资连结保险投资账户 XXXX 年半年信息_