Wednesday, January 11, 2012
Thursday, December 29, 2011
Wednesday, November 2, 2011
Evaluation of Algorithm Performance in ChIP-Seq Peak Detection
제목 그대로 ChIP-Seq 프로그램 performance를 비교한 것인데.. 음 이런 논문이야 말로 짐 회사에서 내기 좋은 주제가 않을까 싶은데.. 아숩다.
Results
overview
3가지 dataset은 NRSF(human neuron-restrictive silencer factor), GABP(growth-associated binding protein), FoxA1(hepatocyte nuclear factor 3a)의 ChIP-Seq데이터.
테스트한 프로그램들 11개의 리스트는 아래와 같고 각 프로그램의 option들은 default를 이용(control data를 이용할 수 있는 프로그램만 선택).
Sensitivity
3개의 dataset에 대해 11개의 program에서 찾는 peak의 수는 차이가 있다.
ChIP-Seq 데이터 분석에서 필요한 peak finding을 위한 프로그램이 31개나 있단다. introduction에 보면 ChIP-Seq 분석 프로그램의 대략적인 알고리즘 개요가 나온다. NGS 특성상 5' 쪽의 tag만 읽기 때문에 생기는 strand-dependent bimodality를 보정하기 위한 방법(paired-end 는 몇개의 프로그램에서만 지원된단다), read가 많이 mapping된 genomic region을 찾는 방법, peak region을 찾기 위해 threshold를 정하는 방법(background signal model 이용:1.manual threshold,2.Poisson or negative binomial model 이용,3.control data 이용), peak의 significance를 정하는 방법에 대한 여러 알고리즘의 간략한 소개가 있다.
일단 여기서는 11개의 peak calling algorithm을 3개의 transcription factor ChIP-Seq 데이터를 가지고 비교한다. 이것이 이 논문의 목적.
Results
overview
3가지 dataset은 NRSF(human neuron-restrictive silencer factor), GABP(growth-associated binding protein), FoxA1(hepatocyte nuclear factor 3a)의 ChIP-Seq데이터.
테스트한 프로그램들 11개의 리스트는 아래와 같고 각 프로그램의 option들은 default를 이용(control data를 이용할 수 있는 프로그램만 선택).
Sensitivity
3개의 dataset에 대해 11개의 program에서 찾는 peak의 수는 차이가 있다.
Wednesday, October 12, 2011
How to Interpret a Genome-wide Association Study
GWAS는 candidate gene analysis(supervised analysis)와 family linkage study, 그리고 HapMAP Project의 성과를 바탕으로 이루어진 것이다.
GWAS는 common disease1, common variant 의 가정하에 있는데 이는 common disease의 유전적인 영향 혹은 원인이 제한된 수의 allelic variant (SNP 혹은 indel 등)에 의한 가정인데 이 제한된 수의 allelic variant 라는 것이 전체 인구에서 1% 혹은 5%이상의 사람들이 가지고 있는 allelic variant를 의미한다.
Overview of GWA Studies
GWAS는 NIH에서 관측된 특성(질병등)의 유전적인 연관성을 찾기 위해 사람 전체 genome에 걸쳐 common genetic variation을 연구 하는 것이라고 정의 되어 있다. genome wide 라는 것의 정확한 기준은 없지만 이 논문에서는 최소한 1,000,000 SNP를 assay 한 연구에 대해서만 언급한다.
GWAS는 크게 4가지 부분으로 이루어진다(PLINK manual에서는 크게 6가지 단계로 나눈다).
1.특정 trait(질병등)를 갖는 집단군(사람들)과 그렇지 않은 집단군을 선택한다.
2.위 단계에서 뽑은 모든 사람을 genotyping 하고 genotyping quality에 대해서 review를 한다
3.2번 단계에서 quality threshold를 넘는 SNP들 중 어느 SNP가 trait와 연관이 있는지 통계적 테스트를 진행한다.
4.3번단계까지 해서 뽑은 genetic variant의 증명단계로 완전 새로운 집단군을 뽑아서 동일한 association test를 진행하거나 아니면 실험적으로 기능적인 효과가 있는지 테스트 한다.
Study Designs Used in GWA
여태껏 가장 많이 사용되었던 GWAS의 design은 case-control design으로 환자군과 정상군의 allel frequency의 비교였다. 이는 가장 간단한 design인데 아래 표와 같이 많은 가정 하에 design 된 연구고 만약 이 가정이 충족되지 않게 된다면 연구의 결과는 상당한 bias가 있게 된다. 그러나 역학적인 디자인의 원리에 충실하게 연구가 디자인 되었다면 case-control은 rare disease에 대한 효과적인 연구가 될 수 있다. 하지만 이게 쉽지 않다는 것.
trio study은 환자와 함께 환자의 부모를 포함한 design이다. 환자인 자식과 그리고 부모의 genotyping을 해서 transmission frequency를 측정한다. 이게 무슨 뜻이냐면 만약 특정 SNP가 질병과 관련이 없다면 부모에서 자식으로의 이전률이 50%일거지만 질병과 관련 있는 SNP는 그 이상일거라는 가정하에 test를 하는 것이다.
또 다른 design은 cohort study이다.
Selection of Study Participants
Genotyping and Quality Control in GWA Studies
Analysis and Presentation of GWA Results
Replication and Functional Studies
Limitations of GWA Studies
Clinical Application of GWA Findings
-----------------------------------reference---------------------------------
1.common disease : mendelian disease이외의 것, mendelian disease 혹은 mendelian disorder 라는 것은 멘델의 유전법칙을 따르는 질환으로 DNA 상의 하나의 mutation으로 인한 질환.
GWAS는 common disease1, common variant 의 가정하에 있는데 이는 common disease의 유전적인 영향 혹은 원인이 제한된 수의 allelic variant (SNP 혹은 indel 등)에 의한 가정인데 이 제한된 수의 allelic variant 라는 것이 전체 인구에서 1% 혹은 5%이상의 사람들이 가지고 있는 allelic variant를 의미한다.
Overview of GWA Studies
GWAS는 NIH에서 관측된 특성(질병등)의 유전적인 연관성을 찾기 위해 사람 전체 genome에 걸쳐 common genetic variation을 연구 하는 것이라고 정의 되어 있다. genome wide 라는 것의 정확한 기준은 없지만 이 논문에서는 최소한 1,000,000 SNP를 assay 한 연구에 대해서만 언급한다.
GWAS는 크게 4가지 부분으로 이루어진다(PLINK manual에서는 크게 6가지 단계로 나눈다).
1.특정 trait(질병등)를 갖는 집단군(사람들)과 그렇지 않은 집단군을 선택한다.
2.위 단계에서 뽑은 모든 사람을 genotyping 하고 genotyping quality에 대해서 review를 한다
3.2번 단계에서 quality threshold를 넘는 SNP들 중 어느 SNP가 trait와 연관이 있는지 통계적 테스트를 진행한다.
4.3번단계까지 해서 뽑은 genetic variant의 증명단계로 완전 새로운 집단군을 뽑아서 동일한 association test를 진행하거나 아니면 실험적으로 기능적인 효과가 있는지 테스트 한다.
Study Designs Used in GWA
여태껏 가장 많이 사용되었던 GWAS의 design은 case-control design으로 환자군과 정상군의 allel frequency의 비교였다. 이는 가장 간단한 design인데 아래 표와 같이 많은 가정 하에 design 된 연구고 만약 이 가정이 충족되지 않게 된다면 연구의 결과는 상당한 bias가 있게 된다. 그러나 역학적인 디자인의 원리에 충실하게 연구가 디자인 되었다면 case-control은 rare disease에 대한 효과적인 연구가 될 수 있다. 하지만 이게 쉽지 않다는 것.
trio study은 환자와 함께 환자의 부모를 포함한 design이다. 환자인 자식과 그리고 부모의 genotyping을 해서 transmission frequency를 측정한다. 이게 무슨 뜻이냐면 만약 특정 SNP가 질병과 관련이 없다면 부모에서 자식으로의 이전률이 50%일거지만 질병과 관련 있는 SNP는 그 이상일거라는 가정하에 test를 하는 것이다.
또 다른 design은 cohort study이다.
Selection of Study Participants
Genotyping and Quality Control in GWA Studies
Analysis and Presentation of GWA Results
Replication and Functional Studies
Limitations of GWA Studies
Clinical Application of GWA Findings
-----------------------------------reference---------------------------------
1.common disease : mendelian disease이외의 것, mendelian disease 혹은 mendelian disorder 라는 것은 멘델의 유전법칙을 따르는 질환으로 DNA 상의 하나의 mutation으로 인한 질환.
Monday, October 10, 2011
Friday, October 7, 2011
Bioconductor: open software development for computational biology and bioinformatics
- primary motivations -
7. Special concerns
CBB에서 생기는 4가지 challenge
-reproducible research
-dynamics of biological annotation
-training
-responding to user needs
- Using Bioconductor (example) -
ALL(Acute lymphocyte leukemia)
- transparency : entire process가 확실하게 노출되어야 한다.
- pursuit of reproducibility : algorithmic work에도 standard가 필요.
- efficiency of development : 기존의 code의 extension과 novice의 발전을 위해 필요.
- seven topics important to establishment of a scientific open source software project -
1. Language selection : 왜 R을 선택 했냐?
-prototyping capabilities : 빠르게 prototype을 만들수 있다. 물론 나중에 더 빠르게 run 할 수 있도록 re-implement도 가능하다.
-packaging protocol : package 형태로 제작, 테스팅, 배포가 가능하다.
-object-oriented programming support : To secure reliable package interoperability
-WWW connectivity : http와 같은 web resource롤 통해 데이터와 package에 접근 가능하며 XML처리하는 package도 있어서 다양한 데이터를 다룰 수(?perceive) 있다.
-statistical simulation and modeling support : R에서 이미 있는 numerical algorithm의 사용이 용이하다.
-visualization support : graphical tool로서의 기능이 좋다.
-support for concurrent computation : parallel computation을 위한 tool들이 있다.
-community : active user and developer communities
-prototyping capabilities : 빠르게 prototype을 만들수 있다. 물론 나중에 더 빠르게 run 할 수 있도록 re-implement도 가능하다.
-packaging protocol : package 형태로 제작, 테스팅, 배포가 가능하다.
-object-oriented programming support : To secure reliable package interoperability
-WWW connectivity : http와 같은 web resource롤 통해 데이터와 package에 접근 가능하며 XML처리하는 package도 있어서 다양한 데이터를 다룰 수(?perceive) 있다.
-statistical simulation and modeling support : R에서 이미 있는 numerical algorithm의 사용이 용이하다.
-visualization support : graphical tool로서의 기능이 좋다.
-support for concurrent computation : parallel computation을 위한 tool들이 있다.
-community : active user and developer communities
2. Infrastructure base
Bioconductor project 에서 첫 2년은 software infrastructure의 투자에 집중한다. 이 infrastructure는 reusable data structure & software 형태로 만든다.
이 software infrastructure concept의 두 예로 Biobase package의 "expreSeq" class와 Bioconductor metadata package 중의 하나인 hgu95av2를 들 수 있다.
expreSeq 은 three-tier architecture 를 용이하게 한다. 뭔말인고 하니 low-level processing software designer는 expreSeq instance만 생성하는데 focusing 하면 되고, 분석가는 low-level processing에 신경쓰지 않고 expresSeq 자료구조만 focusing 해서 분석만 하면 된다.
hgu95av2는.. 잘 모르겠네..
Bioconductor project 에서 첫 2년은 software infrastructure의 투자에 집중한다. 이 infrastructure는 reusable data structure & software 형태로 만든다.
이 software infrastructure concept의 두 예로 Biobase package의 "expreSeq" class와 Bioconductor metadata package 중의 하나인 hgu95av2를 들 수 있다.
expreSeq 은 three-tier architecture 를 용이하게 한다. 뭔말인고 하니 low-level processing software designer는 expreSeq instance만 생성하는데 focusing 하면 되고, 분석가는 low-level processing에 신경쓰지 않고 expresSeq 자료구조만 focusing 해서 분석만 하면 된다.
hgu95av2는.. 잘 모르겠네..
3. Design strategies and commitents
-designing by contract
-object-oriented programming
-modularization
-multiscale and executable documentation
-automated software distribution
-designing by contract
-object-oriented programming
-modularization
-multiscale and executable documentation
-automated software distribution
4. Distributed development and recruitment of developers
distrubted development의 강조. CVS를 통한 같은 component에 대해 여러 developer가 개발에 참여하게 됨으로 다양한 viewpoint와 experience가 project에 속하게 된다. 한 사람의 개발자에 의한 code의 변화가 다른 코드를 망가지게 하지 않는 것을 원칙으로 한다. 이는 사실 packaging화로 가능하다. 그리고 이 R package의 규격화된 testing system을 제공함으로 인해 안정적인 개발이 가능케 한다.
distrubted development의 강조. CVS를 통한 같은 component에 대해 여러 developer가 개발에 참여하게 됨으로 다양한 viewpoint와 experience가 project에 속하게 된다. 한 사람의 개발자에 의한 code의 변화가 다른 코드를 망가지게 하지 않는 것을 원칙으로 한다. 이는 사실 packaging화로 가능하다. 그리고 이 R package의 규격화된 testing system을 제공함으로 인해 안정적인 개발이 가능케 한다.
5. Reuse of exogenous resources
다른 project의 software를 adapting 하는데 있어서의 3가지 쟁점
-가능하면 re-implementation 하지말고 있는거 갖다 쓰자.
-CBB(computational biology & bioinformatics)는 다양한 분야를 아우르기 때문에 다른 많은 프로젝트와의 공동의 노력이 필요하다. 그렇기 때문에 다른 언어나 시스템에서 쓰여진 데이터나 알고리즘을 사용하기 위한 구조화된 패러다임이 필요하다.
-standardization and reuse of existing tools
다른 project의 software를 adapting 하는데 있어서의 3가지 쟁점
-가능하면 re-implementation 하지말고 있는거 갖다 쓰자.
-CBB(computational biology & bioinformatics)는 다양한 분야를 아우르기 때문에 다른 많은 프로젝트와의 공동의 노력이 필요하다. 그렇기 때문에 다른 언어나 시스템에서 쓰여진 데이터나 알고리즘을 사용하기 위한 구조화된 패러다임이 필요하다.
-standardization and reuse of existing tools
6. Publication and licensing of code
7. Special concerns
CBB에서 생기는 4가지 challenge
-reproducible research
-dynamics of biological annotation
-training
-responding to user needs
- Using Bioconductor (example) -
ALL(Acute lymphocyte leukemia)
Friday, September 16, 2011
Enrichment of differentially methlyated regions with MethylMiner fractionation and deep sequencing with the SOLID System
invitrogen 에서 MethylMiner 라는 kit의 application note로 publish 한 것(여기). 이걸 본 이유는 elution의 농도에 따른 bias를 어떻게 해야 하는가에 대한 의문에서인데..
생각보다 MBD-Seq의 장점을 많이 알 수 있게 되서 posting 한다(물론 invitrogen에서 자기네 kit 홍보성으로 만든 것이기에 공정성은 좀 떨어지지만 그래도..).
MBD-Seq의 장점 (MeDIP-Seq 대비)
antibody로 DNA methylation을 detecting 할려면 DNA 를 denature 시켜야 한단다. antibody가 mCpG에 fully access해야만 하기에 denaturation 시키는 과정이 들어가게 되고 이는 실험적인 소요가 요구된다. 음.. 전혀 생각 못했던거다. 워낙에 논문들 대세가 MeDIP-Seq 이라서 그냥 생각없이 antibody를 이용한 precipitation이 좋을 줄 알았는데..
글고 여기서 보여지금 그림 하나.
그리고 반드시 elution 농도를 섞어에 하는게 아닌가 싶다. 이 application note에서는 500mM과 1M 만을 비교 했는데.. 언뜻 보면 1M 이 좋아보이는가 싶지만 마지막 그림에서 보이듯이 500mM에서만 유독 enrichment 되는 부위가 있다. 곧 모든 elution 농도로다가 뽑아낼 수 있는 DNA는 전부 뽑는게 가장 좋은 방법이 아닐까 한다만... 모르겠다.
생각보다 MBD-Seq의 장점을 많이 알 수 있게 되서 posting 한다(물론 invitrogen에서 자기네 kit 홍보성으로 만든 것이기에 공정성은 좀 떨어지지만 그래도..).
MBD-Seq의 장점 (MeDIP-Seq 대비)
antibody로 DNA methylation을 detecting 할려면 DNA 를 denature 시켜야 한단다. antibody가 mCpG에 fully access해야만 하기에 denaturation 시키는 과정이 들어가게 되고 이는 실험적인 소요가 요구된다. 음.. 전혀 생각 못했던거다. 워낙에 논문들 대세가 MeDIP-Seq 이라서 그냥 생각없이 antibody를 이용한 precipitation이 좋을 줄 알았는데..
글고 여기서 보여지금 그림 하나.
위 그림이 MeDIP 처리 한거랑 MethylMiner 처리한거랑 dye를 붙여서 chip으로 찍었을때 어느 것이 enrichment되나를 본건데 MethylMiner를 이용한것이 훨씬 많은 array hit에서 enrichment 가 되어 있음을 알수 있다. 단순하게 생각한다면 MethylMiner가 훨씬 sensitive 한것처럼 보인다.
elution salt 농도는 어쩌라는 거냐?
음.. 여기 있는거 가지고 굳이 결론을 내려보자면(엄연히 내 생각이고.. 실험에 대한 이해 부족으로 잘못된 결론일 수 있다는 가능성은 농후하다)..
이란 언급에서 500mM이랑 1M만으로 sequentially elution 했음에도 거의 대부분의 captured DNA가 elution 됨을 알수 있다. 그렇기 때문에 gradient로다가 elution을 하게 되면 거의 모든 captured DNA 가 elution 될 것이므로 gradient에 따른 특정 CpG density 영역의 read들만 elution 되는 일은 없을 것으로 생각되어 진다."Greater than 90% of the captured DNA is sequentially eluted with 500 mM and 1M NaCl"
그리고 반드시 elution 농도를 섞어에 하는게 아닌가 싶다. 이 application note에서는 500mM과 1M 만을 비교 했는데.. 언뜻 보면 1M 이 좋아보이는가 싶지만 마지막 그림에서 보이듯이 500mM에서만 유독 enrichment 되는 부위가 있다. 곧 모든 elution 농도로다가 뽑아낼 수 있는 DNA는 전부 뽑는게 가장 좋은 방법이 아닐까 한다만... 모르겠다.
근데 또 이 논문을 보긴 봐야 할거 같은데.. 언뜻 method 만봐서는 잘 이해가 안가긴 하지만..
Subscribe to:
Posts (Atom)







