Decoding gene expression regulation through motif discovery and classification
Our rough guess is there are 30,000 words in this book.
At a pace averaging 250 words per minute, this book will take 2 hours and 0 minutes to read. With a half hour per day, this will take 4 days to read.
How long will it take you?
This book will take an estimated to read at a reading speed averaging words per minute. With 30 minutes per day, this will take to read.
Enter your reading speedYou can take one of our WPM reading speed tests to find your reading speed.
Create a free account to track your reading progress, build your reading list, and set reading goals.
Word Count
30,000 words, Guess
Page Count
120 pages
Identifiers
- OCLC Control Number477172656
- Open LibraryOL45212590M
Description
Biological systems are complex machineries with numerous components interacting with each other. Through the regulation of gene expression, the systems work differently at different conditions. The regulatory rules are by and large determined by DNA, as it is the most important inheritable substance. Thus, it is interesting to infer these rules by building connections between DNA sequences and gene expression. Modern high-throughput technologies are able to provide us with massive amounts of data related to sequence features and gene expression. However, the scale of the data also brings the challenges of variable selection and computation efficiency. This dissertation presents several biology problems which involve DNA motif discovery and gene regulatory rule inference through the development of graphical models and variable selection techniques. The first chapter introduces some basic biological concepts of DNA sequence analysis and regulatory network construction in computational biology. The second chapter discusses motif discovery problem, including its current status and challenges, with a real data application of motif finding for protein abrB in Bacillus subtilis, using a specially designed protein binding microarray data. In the third chapter, we present the problem of predicting gene expression using DNA sequences. Sequence features such as motif scores are used as predictors. A Bayesian variable selection scheme is designed to select motifs which are most related to the expression of target genes, and also discover the interaction or synergic effect among them. This method is further extended into a general classifier, called selective partially augmented naïve Bayes (SPAN). The fourth chapter compares this classifier and its variant C-SPAN to several state-of-the-art classifiers, with applications in several real and simulated datasets. SPAN is a very fast classifier, and is shown to have an intrinsic connection with logistic regression It is able to fit a logistic regression model with large number of covariates, achieving both variable selection and interaction detection at the same time.
Subjects
Reader Reviews
No reviews yet for this book.
Be the first to share your thoughts!