Helping Sustained Learning

Right since our childhood, we have been force fed by our education system to mug up all that we can. No one knows why. The teacher's don't. The parents don't. The children for sure don't.

Given the fact that most of the human knowledge is already on the internet (read Wikipedia), don't you think we need to break the rules now? Do any of us grown up adults remember anything more than how to find stuff from the Net and use it logically? I, for one, sure do. So, do we really need our children to go through the excruciatingly painful education system?

Aside: Wow!! I really like to ask questions.

Aside II: Can I provide answers? Sure why not!! What I propose is something like a secondary brain. Our vault of memory on the cloud.

Like it or not Sherlock Holmes was right. We do have a limit to the Petabyte HDD we carry in out heads. We can only store so much information in our head. Why bug it up with all the unnecessary information children get from the subjects they are not interested in?

Why not give them a Single-point-of-access-device for all their information needs? Whatever they read or see can be stored on the cloud so they don't need to remember anything (if they wish to). All that is required of them is that they process that information quickly and logically.

We already have systems in place for all that I propose. All we need is convergence and social acceptability.

Spare a thought for it!!

The Social Agenda

Do Social networks really work?

Yesterday while talking to my boss (BTW he is the only one I still consider to be my boss, apart from my wife, as he is the only one I listened to at work ;)), he came up with an interesting idea. People have not changed since centuries. They have always been drifting apart. From whole villages living in one cave to rising divorce rates, we have come a long way from social living. People are becoming more and more individualistic.

Now, back to the main topic. Do social networks work? For the past decade, we have seen numerous social networks coming and then fizzling out. MySpace, Friendster, Orkut and now Facebook. All have seen the same story. People get excited when they are launched. Call all their friends to join them. They form groups and communities. They chat all night long. But then after that what? They move on to the next big thing... Twitter.

For the fear of being called a hypocrite, I would like to admit that I have also been a part of all the bandwagons. Be it Orkut, MySpace, Blogs, Facebook or Twitter. You can find my profiles on all of them. But then I become bored like the rest of you and move on.

Now that Google is coming onto the scene with a stronger product (than Orkut) in Google +, I am not sure whether it is too late to make any dent in the SNSs' fortunes.

There are three things and only three things that really sell in this world. Sex, knowledge and food. That is why you see all the Porn sites being so popular and Google making tonnes of money. Online economies around these three will always remain popular and viable. The rest as they say will become history.

BTW with the much hyped IPOs of all the SNSes, I seriously feel there is another dot com burst brewing. So cash in your stocks and come back to the real world!!

Cloud Computing and Television

This post the use of Cloud Computing to provide true ‘On Demand Entertainment’ and make television a Unified Entertainment Device. This paper uses some interesting use-cases for Televisions which can be possible by using the power of Cloud Computing to illustrate this. These applications can transform televisions to be a complete entertainment platform.

Introduction

Televisions have come a far way from their monochromatic ancestors of the 1930's to become the modern age's ultimate home entertainers. Demands and expectations of the consumers have evolved tremendously to put a strain on the traditional mediums of entertainment, i.e. audio and video. Moreover, the advent of internet and has created another dimension in home entertainment – lets call it "On Demand Entertainment" (ODE). True ODE means consumers will be able to watch, listen, play or read whatever they want whenever they want.

The consumers of Televisions now have options of going on the Internet and search for ODE including (but not restricted to) games, news, video and audio. Internet giants like Amazon, Hulu, Netflix and Youtube have started cutting into Television industry's profits and have become a major force to reckon with in home entertainment segment.

In a parallel development, consumers now want seamless integration and convergence between the new and the old media of entertainment. This expectation has led to innovation in Television industry in the form of IPTV, satellite TV and internet enabled TVs. These technologies strive to provide ODE as well as try to fulfill the demand for a unified entertainment device. However, true ODE is still a distant dream because of the strain it puts on the storage and computation power of the back-end data centers.

Cloud Computing or internet based computing, which provides on demand storage and compute power to be billed in a pay-per-use basis, comes as a perfect strategic fit to solve the puzzle of ODE. Cloud Computing can provide a solution to the issue of huge requirements in compute and storage to provide true ODE.

This post describes how Cloud Computing can be used to deliver true On Demand Entertainment, using some specific use-cases of:
  • On Demand Gaming
  • Ubiquitous Media Playback
  • Online Personal Media Store
The Workings!!

Entertainment today includes much more than the traditional media of books, Television and Radio. As we discussed earlier, ODE has become a major expectation now-a-days. Also, people are now looking for a single device which can take care of all their entertainment needs. Televisions are facing some serious competition in this race for a unified entertainment device from hand held gadgets and Internet. Televisions need help of modern technology to break to the fore-front of this race. Cloud Computing is one such technology which can tremendously help the Television industry.

According to Wikipedia:

Cloud Computing is Internet-based computing, whereby shared resources, software, and information are provided to computers and other devices on demand, like the electricity grid.

This on-demand Cloud of servers, generally called the cloud, can provide for the huge hardware requirements for some interesting use cases in Television industry:

On Demand Gaming:

Games are very compute intensive applications. So much so that they have dedicated platforms built for serious gamers. Televisions too have in-built games but not of the class of "Core Games". This is because core games require huge compute power that Televisions can't provide.

Now that Televisions have become internet enabled, we can use the compute power of the Cloud to do the computation at the back-end. We can push the gaming consoles on the cloud. All the user interactions can be pushed onto the Cloud; the cloud will compute based on the game rules and send back the results for the Television to display.

This can be a disruptive product in the gaming industry as this will give rise to true multi player games, where players join and leave the games as and when they will. Anyone with an internet enabled Television set can join the game as dependency on expensive gaming consoles will end.

Ubiquitous Media Playback:

Another interesting area of application of Cloud Computing in entertainment industry is of Ubiquitous Media Playback.

Let's take an example:

User is watching a movie on his Television when he suddenly gets an urgent call to go somewhere. The movie is at a very interesting phase and he doesn't want to miss it. He can simply activate his “Ubiquitous Media Playback” feature with the push of a button and the movie starts playing back on his hand held gadget.

For this to become reality, all that is needed is that both his hand held gadget and Television have internet access. The Television starts uploading the movie (from the point where it was stopped) to the cloud; the cloud converts the movie to be fit to playback on his hand held gadget and streams it to the gadget. The gadget resumes playing the movie from the same point.

Thus, a simple application of the power of cloud computing can enhance the viewer experience manifolds.

Online Personal Media Store:

People now try to keep all their media in digital format. They store it on hard disks, CDs, DVDs and BDs. But it still forms a bulky collection with the chance of loosing their data always lingering on top of their mind. What if they had a Personal media Store on the Cloud? What if they can use their Televisions to store all they want on the cloud?

This can be possible. Televisions connected to the Internet can be used to dump all the personal media on the cloud. The Cloud can then sort and organise the media under various categories and make a searchable index of all the user content. This data can then be customized according to the display and other capabilities of all the user‟s devices. The user can then access this library of data using any device he wants.

This can be a pay per use service which can be very easily commercialized.

These are just a few use cases which illustrate the power of Cloud Computing in home entertainment. If we take a sweeping look at the whole home entertainment landscape, there would be thousands of applications of this tremendously powerful technology.

Conclusion

We know that technology is pushing the edges day in and day out. Televisions need to adopt these new technologies rapidly, in order to provide a better experience to the consumers. Cloud Computing is one such technology which has the power of revolutionizing the way entertainment is served. As this paper suggests, the use of this new technology with different applications can make the true ODE provider and a unified entertainment device.

Automatic program recommendation using Facial Expression Recognition

More than often people get bored of the program they are currently watching. However, they dread scanning through hundreds of channels to find out if anything interesting is coming. They rather end up switching off the TV. This idea proposes a solution to these situations: The TV suggests the viewer what’s interesting on the air.

The TV, using a camera attached to it, first recognizes the viewer using “Face Recognition” techniques. It then relays the statistics about the time the viewer views a particular program and his facial expressions (happy, sad, interested, disinterested, etc) while watching it to a central hub.

The central hub (which may be on the cloud or in a self managed datacenter) then collates this data with the ontology of the program. [This ontology is built by data mining and supervised machine learning using a distributed cluster of computers.] Thus, over time, the central hub builds a profile of the viewer according to the viewer’s viewing pattern.

In the meanwhile, the TV also recognizes the facial expression of the viewer and looks for the signs of boredom (drowsy eyes, long frowns, vacant stares, etc). As soon as it recognizes a “disinterested pattern”, it calls a web-service running on the central hub.

This web-service searches the ontology of programs currently on air and sorts it according to the viewer’s profile. Then it predicts a list of programs which will most likely interest the viewer and sends it back to the TV.

The TV then suggests the viewer about the programs (s)he can watch if he is not interested in the current one. If the viewer changes the program then the TV sends the information to the central hub about the program he switched to. The Central hub adds this information to the profile of the viewer.

In case of conflicts, i.e. more than one viewer, the central hub decides on the predicted list of interesting programs using a priority list of the viewers based on previous history as to whose choice prevailed last time such a conflict arose.

This figure describes the architecture in a simple yet concise manner:


Note: If anyone wants to use this idea or wants further clarification then please contact me @ paritosh (dot) gunjan (at) gmail (dot) com.

Install Hadoop on a Single Machine

In order to learn any new technology, the best way IMHO is to install it on your local system. You can then explore it and experiment on it at your convenience. It is the route I always take.

I started with installing Hadoop on a single node, i.e. my machine. It was a tricky task with most of the tutorials making many assumptions. Now that I have completed the install, I can safely say that those were simple assumptions and that anyone familiar with linux is deemed to understand them. Then I decided to write an installation instruction for the dummies.

Here is the most comprehensive documentation of "How to install Hadoop on your local system". Please let me know if I have missed anything.


Prerequisites


1. Linux

The first and foremost requirement is to get a PC with Linux installed on it. I used a machine with Ubuntu 9.10 installed on it. You can also work with Windows, as Hadoop is purely java based and it will work with any OS that can run JVM (which in turn implies pretty much all the modern OS's)

2. Sun Java6

Install the Sun Java6 on your Linux machine using:
$ sudo apt-get install sun-java6-bin sun-java6-jre sun-java6-jdk

3. Create a new user "hadoop"

Create a new user hadoop (though it is not required, it is recommended in order
to to separate the Hadoop installation from other software applications and user
accounts running on the same machine by having a dedicated user for hadoop).
Use the following commands:
$ sudo addgroup hadoop
$ sudo useradd -d /home/hadoop -m hadoop -g hadoop

4. Configure SSH

Install OpenSSH­Server on your system:
$ sudo apt­-get install openssh­-server

Then generate an SSH key for the hadoop user. As the hadoop user do the
following:
$ ssh-­keygen -­t rsa ­-P ""
Generating public/private rsa key pair.
Enter file in which to save the key (/home/hadoop/.ssh/id_rsa):
Created directory '/home/hadoop/.ssh'.
Your identification has been saved in /home/hadoop/.ssh/id_rsa.
Your public key has been saved in /home/hadoop/.ssh/id_rsa.pub.
The key fingerprint is:
1a:38:cd:0c:92:f9:8b:33:f3:a9:8e:dd:41:68:04:dc hadoop@paritosh­desktop
The key's randomart image is:
+­­[ RSA 2048]­­­­+
|o . |
| o E |
| = . |
| . + * |
| o = = S |
| . o o o |
| = o . |
| o * o |
|..+.+ |
+­­­­­­­­­­­­­­­­­+
$


Then enable SSH access to your local machine with this newly created key:
$cat $HOME/.ssh/id_rsa.pub >> $HOME/.ssh/authorized_keys

Test your SSH connection:
$ ssh localhost
The authenticity of host 'localhost (::1)' can't be established.
RSA key fingerprint is 1e:be:bb:db:71:25:e2:d5:b0:a9:87:9a:2c:43:e3:ae.
Are you sure you want to continue connecting (yes/no)? yes
Warning: Permanently added 'localhost' (RSA) to the list of known hosts.
Linux paritosh­desktop 2.6.31­20­generic #58­Ubuntu SMP Fri Mar 12 05:23:09 UTC
2010 i686
$

Now that the prerequisites are complete, lets go ahead with the Hadoop
installation.

Install Hadoop from Cloudera

1. Add repository

Create a new file /etc/apt/sources.list.d/cloudera.list with the following
contents, taking care to replace DISTRO with the name of your distribution (find
out by running lsb_release -c):
deb http://archive.cloudera.com/debian DISTRO­-cdh3 contrib
deb­src http://archive.cloudera.com/debian DISTRO­-cdh3 contrib
2. Add repository key. (optional)

Add the Cloudera Public GPG Key to your repository by executing the following
command:
$ curl -­s http://archive.cloudera.com/debian/archive.key | sudo apt­-key add ­-
OK
$

This allows you to verify that you are downloading genuine packages.

Note: You may need to install curl:
$ sudo apt­-get install curl

3. Update APT package index.

Simply run:
$ sudo apt­-get update

4. Find and install packages.

You may now find and install packages from the Cloudera repository using your favorite APT package manager (e.g apt­-get, aptitude, or dselect). For example:
$ apt-­cache search hadoop
$ sudo apt­-get install hadoop

Setting up a Hadoop Cluster

Here we will try to setup a Hadoop Cluster on a single node.

1. Configuration

Copy the hadoop­0.20 directory to the hadoop home folder.
$ cd /usr/lib/
$ cp ­-Rf hadoop­0.20 /home/hadoop/
Also, add the following to your .bashrc and .profile
# Hadoop home dir declaration
HADOOP_HOME=/home/hadoop/hadoop­0.20
export HADOOP_HOME
Change the following in different configuration files in the /$HADOOP_HOME/conf dir:

1.1 hadoop­env.sh

Change the Java home, depending on where your java is installed:
# The java implementation to use. Required.
export JAVA_HOME=/usr/bin/java

1.2 core-­site.xml

Change your core­-site.xml to reflect the following:

1.3 mapred-site.xml

Change your mapred-­site.xml to reflect the following:

1.4 hdfs-site.xml

Change your hdfs­-site.xml to reflect the following:

2. Format the NameNode

To format the Hadoop Distributed filesystem (which simply initializes the directory specified by the dfs.name.dir variable), run the command:
$ /hadoop/bin/hadoop namenode -format

Running Hadoop

To start hadoop, run the start­all.sh from the
/$HADOOP_HOME/bin/ directory.

$ ./start-all.sh

starting namenode, logging to /home/hadoop/hadoop-0.20/bin/../logs/hadoop-hadoop-namenode-paritosh-desktop.out

localhost: starting datanode, logging to /home/hadoop/hadoop-0.20/bin/../logs/hadoop-hadoop-datanode-paritosh-desktop.out

localhost: starting secondarynamenode, logging to /home/hadoop/hadoop-0.20/bin/../logs/hadoop-hadoop-secondarynamenode-paritosh-desktop.out

starting jobtracker, logging to /home/hadoop/hadoop-0.20/bin/../logs/hadoop-hadoop-jobtracker-paritosh-desktop.out

localhost: starting tasktracker, logging to /home/hadoop/hadoop-0.20/bin/../logs/hadoop-hadoop-tasktracker-paritosh-desktop.out

$


To check whether all the processes are running fine, run the following:

$ jps

17736 TaskTracker

17602 JobTracker

17235 NameNode

17533 SecondaryNameNode

17381 DataNode

17804 Jps

$

Hadoop... Let's welcome the Elephant!!

While working with my crawler, I came to a conclusion that I can't work with my poor old system for processing such a large amount of data. I needed some powerful machine and huge amount of storage to do it. Unfortunately (or fortunately for me) I don't have that kind of moolah to invest in a enterprise class server and a huge SAN Storage. What began as a pet project was growing out to become a major logistics headache!!

Then I read something about a new open source entrant into the Distributed Computing space called Hadoop. Actually its not a new entrant (has been in development since 2004, when the Google's MapReduce algorithm paper was published). Yahoo has been using it for similar purposes of creating page indexes for Yahoo Web Search. Also, Apache Mahout, a machine learning project from Apache Foundation, uses Hadoop as its compute horsepower.

Suddenly, I knew Hadoop is the way to go. It uses commodity PCs (#), gives Petabyes of storage and the power of Distributed Computing. And the best par about it is that it is a FOSS.

You can read more about Hadoop from the following places:

1. Apache Hadoop site.
2. Yahoo Hadoop Developer Network.
3. Cloudera Site.

My next few posts would elaborate more on Hadoop's working.

Note:

# Commodity PC doesn't imply cheap PCs, a typical choice of machine for running a Hadoop datanode and tasktracker in late 2008 would have the following specifications:

  • Processor: 2 quad-core Intel Xeon 2.0GHz CPUs
  • Memory: 8 GB ECC RAM
  • Storage: 41 TB SATA disks
  • Network: Gigabit Ethernet

Machine Learning : The Crawler

Building the crawler was the easiest part of this project.

All this crawler does is take a seed blog (my blog) URL, run through all the links in its front page and store the ones that look like a blogpost URL. I assume that all the blogs in this world are linked to on at least one other blog. Thus all of them will get indexed in this world if the spider is given enough time ad memory.

This is the code for the crawler. Its in Python and is quite easy. Please run through it and let me know if there is any other way to optimise it further:

import sys
import re
import urllib2
import urlparse
from pysqlite2 import dbapi2 as sqlite

conn = sqlite.connect('/home/spider/blogSearch.db')
cur = conn.cursor()

tocrawltpl = cur.execute('SELECT * FROM blogList where key=1')
for row in tocrawltpl:
tocrawl = set([row[1]])

linkregex = re.compile("")

while 1:

        try:
                crawling = tocrawl.pop()
        except KeyError:
                raise StopIteration

        url = urlparse.urlparse(crawling)

        try:
                response = urllib2.urlopen(crawling)
        except:
                continue

        msg = response.read()
        links = linkregex.findall(msg)

        for link in (links.pop(0) for _ in xrange(len(links))):
                if link.endswith('.blogspot.com/'):
                        if link.startswith('/'):
                                link = 'http://' + url[1] + link
                        elif link.startswith('#'):
                                link = 'http://' + url[1] + url[2] + link
                        elif not link.startswith('http'):
                                link = 'http://' + url[1] + '/' + link

                        select_query='SELECT * FROM blogList where url="%s"' %link
                        crawllist = cur.execute(select_query)
                        flag=1
                        for row in crawllist:
                                flag=0

                        if flag:
                                tocrawl.add(link)
                                insert_query='INSERT INTO blogList (url) VALUES ("%s")' %link
                                cur.execute(insert_query)
                                conn.commit()

Machine Learning : Emotional Intelligent Search Engines

When we think of machines being emotionally intelligent, we tend to see a vast problem with no seemingly plausible solution. However, as I stated in my post about Machine Learning and Semantic Web, it can be broken down into smaller, manageable problems. And I intend to do that.

Here, in this post I present the introduction to a small step towards this solution. This solution, if implemented properly (which I hope it will be), will allow a machine to return emotionally intelligent search results.

I assume that anyone reading this article knows the way a search engine works, the Python programming language and how databases work.

Statement of Problem :

"Build a emotional intelligent search engine which does the following tasks:

  1. Crawls the web for blogposts and downloads them.
  2. Indexes them according to their emotional quotient.
  3. Returns results according to the abstract questions asked by the end user. "
Proposed solution :

The proposed solution POC consists of :
  1. A crawler which crawls the blogosphere collecting links of all the blogs out there. I intend to target only the blogs on blogspot. These links are sanitised and stored in a database. Then another crawler cum 'blog-getter' looks into this database, picks up each blog one by one and then downloads all the posts in them. It stores the posts in flat file format and stores its metadata in a database.
  2. A lexicon builder parses these posts one by one and builds upon its vocabulary of emotional words (fundamentally all the adjectives in the posts) with the help of a human tutor.
  3. An indexer then parses all these posts, cross referencing them with its lexicon (built in the previous step) and then stores the parsed words into a db. It stores nouns and adjectives in separate tables referencing them to the metadata table (built in the first step).
  4. When the end user enters a search query, suppose "a nice Chinese restaurant", the query parser first finds all posts about 'Chinese restaurant'. From these results it finds the ones that are nice (according to its logic 'nice' means where people were happy and used 'happy' adjectives when they described such restaurants in their posts). It then displays the results in descending order of 'niceness'.
Implementation

I sincerely believe that this is an absolutely doable project, and I have made some progress with building the lexicon. As this is an ongoing project, I will be posting my experiences and the codes as and when I complete testing them.

The coding will be in Python. Why? Because I am in love with the language. Period.

The database, as of now is SQLLite, but I am planning to convert it into a NoSQL as soon as I get a firm grasp on the technology.

Platform right now is a single Linux box. I know it will require a lot more computing power, storage and memory when deployed. I am planning to port this to Hadoop when it comes to deployment stages. But for now, this remains my baby and will have its place in my desktop :).

Supervised Machine Learning

Introduction

When I was born, I was ignorant but very curious. I always asked
questions : "What? Where? Why? How?" And my parents guided me through
this jungle of knowledge, telling me about things that really mattered
and leaving what was complicated to the future.

As I grew up, I started learning about the very complex things by
cross referencing it with my already existent knowledge. If the new
incidents were not in my knowledge database, I queried the books,
articles and lately Google. Thus from an ignorant child with an empty
knowledge database, I grew up to become a knowledgeable person.

Let me take an example of how I learned English.

I am from India, English is neither my nor my parents' first language.
However, they really wanted me to learn how to read, write and speak
in English. They couldn't teach me by changing their mother tongue.
What they did, instead, was that they taught me the alphabets, a few
easy words and the basic grammar constructs. Then they gave me a
dictionary and a grammar book. I referenced new words that I saw with
the dictionary and new sentence formations with the grammar book. If I
still couldn't understand what was written, I asked my teacher.

Now when I come across anything written in English, I cross reference
it with the vocabulary and grammar that I have learned all these
years. If something new comes up, I check it up online or with my
peers. Thus I build upon my knowledge about English in an ongoing,
continuous process.

Implementation in Machine Learning

Now, building on this very basic pattern how any human child learns
about this physical world, we can model this learning process for the
machines. This is a brief outline of the idea (here I take an example
of the English Language, but this is equally applicable to any other
natural language or any other real world entity):

"In the same way as I learnt about the world, Machines can also be
taught. In the same way that we teach a child how to read and
understand the English Language, a computer can be taught to recognize
patterns of English alphabets as words and can reference its meaning
with its dictionary (which is built by human assistance). The machine
can gradually build up its vocabulary by reading more and finding the
meanings of new words and sentence constructs online. If it finds
something which is complex enough, it can ask its human tutor its
meaning.

Given the amount of cheap storage, memory and computing power we have
in our hands today, I think it is a perfectly doable thing. I just
need time and resources to do it."

Value of this research

Why is it important for a machine to learn a natural language? The
answer is simple, if machine learns a natural language, it can mine
the enormous amount of data on the internet to learn about any other
discipline. We will provide the world with the perfect student. It can
learn financial modelling from a human tutor and apply it to the
streaming financial data from the internet to predict future trends.
Maybe even produce its own models. It can learn biotechnology from its
human tutor and mine the huge available pathological data on the
internet to produce new medicines. The application possibilities seem
to be bound only by human imaginations.

Current Status of my Research

Right now, I am working on a search engine which crawls the
blogosphere and indexes blogposts according to their emotional
quotient. Its a small step in the final scheme of things but this is
what I can do with the time and resources that I have (I have a '9 to
6' day job as a software engineer). I would love to devote more time
to this idea, given the mundane activities (like earning my bread and
butter) is taken care of :-).

I am not sure if this idea has been implemented previously anywhere.
Also, I acknowledge it is a very simple solution to a seemingly very
complex problem. However, I feel that this is the most suitable
solution. If we humans can do it then so can the machines.

Microsoft Interops Day

Well, well, well... I committed the crime. An open source developer attending a Microsoft event is definitely a crime. But then they promised it was going to be an event for PHP development on Windows. I took the bait, hook, line and sinker. Damn!!

The interoperability day turned out to be a cheap marketing gimmick, showcasing what you can do with IIS and Silverlight. This is a session by session breakup of what the Microsoft guys showed us:

Session 1:
Promised : Build an end-to-end, authentication and authorization model for your existing PHP application in less than 30 mins.

Delivered : An existing PHP application was taken, authentication provided by removing anonymous access from the root directory, and authorization provided by adding a login page. Quite a good demonstration , I would admit. With a few clicks of mouse, we could achieve something which takes a few lines of code.

My Take : To a hand coder like me, who needs to have control on all the stuff going around me, I don't think I will ever use this stuff.

Session 2:
Promised : Access your data exposed from your PHP app, through a RESTful service interface.

Delivered : This is where I dozed off. The demo was to show how you can easily configure a webservice for a SQL Server DB, plugin your CS code to harvest this code and show it through your browser.

My Take : Now, I may be snoring while they showed how the PHP application comes into the picture, but I have no memory of any PHP code being shown in this session actually. It was all about how to use standard plug n play snippets in Visual Studio, ADO.NET and SQL Server.

Session 3:
Promised :
Configuring Windows to run your PHP applications.

Delivered : They showed us some so-called cool GUI configuration techniques for porting .htaccess configurations.

My Take : Now why would I do that? When I have a single .htaccess file to control all my apache configuration, why would I monkey around with the superfluous click-and-you-are-done controls in IIS? And the best part of the presentation was that the guy showing this all leads the PHP UG in Bangalore. What a sham!

Session 4:
Promised :
Discover the web designer in you, to build pure CSS based PHP websites.

Delivered : A couple of CSS tips and tricks downloaded from the net. And a huge demo of how Silverlight can Light up your web apps.

My Take : No discussion what-so-ever for using CSS for PHP templates!! Cummon!! Atleast stick to what was promised in your mailer. But then MS is never known to stick to its promises, is it?

Overall, I think that if they had called us in for a demo of what MS is doing for the Web Developers, then I would have definitely loved the sessions. I have never worked on MS stuff for web apps development and I would have loved to have a kick start. But I went there with a differend picture in mind, and was deeply disappointed by the way in which the MS guys used the name of PHP for promoting their products.

N.B. I am currently working on AWS using Python. I find it a compelling cloud computing platform and will be starting a series of posts on it. Anyone who loves cloud computing, hold on to your keyboards!! ;)

Blogadda looses my data!!

Dear Blogger Friend,

Thank you for visiting BlogAdda.com and submitting your blog.


We had an issue with our host yesterday due to which information about your blog was inadvertently affected. Your blog URL is in the records and you might have to update the rest of the information.

Kindly login at BlogAdda (
http://www.blogadda.com/login
) and click on 'My Account' link on the top. You'll see the title of your blog linked. Clicking the title will expand into a few options and you can click on 'Manage' to update details of your blog.

Our apologies for the inconvenience caused and we would appreciate if you can update the blog details asap.

We would like to assure you that we take utmost care of the data and this would not happen again. We thank you for your co-operation.


Regards

Administration Team

BlogAdda.com

Twitter:
http://www.twitter.com/blogadda
Facebook:
http://www.facebook.com/blogadda

What happenes when you get such a mail? I get pissed off. Within a minute of getting this mail, I shot back my reply:

No need to be sorry!! I have got tonnes of work other than to re-enter my blog details on your stupid servers which can't even store this tiny winy bit of information properly.

I am logging out of blogadda. I will make sure no one in my friendlist joins or advises anyone to join such a site which doesn't value its user's data.

I know blogadda is a free blog aggregator service and they have no responsibility what-so-ever for my data. But the question remains... how the hell can they manage to loose it?

There are many ways to loose data:

  1. Hacking : Someone hacked into their database and snooped off with the data. In that case I my email ID is in their hands and at the very least, I can expect to get thousands of spam emails which may or maynot have adware and malwares attached to them.
  2. Backup failure: Some bozo missed out taping my data. If they have a backup policy. Which I doubt hey have.
  3. Natural Disaster: The data centre was flooded and thus my data was lost.
  4. Terrorist attack: The black masked terrorists came and fled off with the tapes of my data. They also deleted my entries from the database. On the second thoughts, is my second name Bush? Definitely not.

After going through all such possibilities, all I can say is that I am mesmerised by this kinda act of terrorism on my individuality. Damn it!

A Brief History of JavaScript : Part I

"Javascript is not Java", it is the common phrase we always read whenever we read any book on JavaScript. But then why keep similar names? Its like saying the Red Indians are not from India. Huh!

Before we get into the debates about the mistakes driven by the marketing at Sun Microsystems and Netscape, which indeed has plagued web designers for years and would do so for years to come, lets have a look at the twisted history of JavaScript.

Back in the 1990s, the web was dominated by static websites.

The first graphical web browser , Mosaic, was launched by in December 1995 by NCSA. Its main achievement was that it was also the first browser to display images inline with text instead of displaying images in a separate window. Though the web pages still remained static but more color could be added to them.

This small advancement in the rendering of the web pages led to the frantic rush in development and selling of advanced browsers. People wanted the guys who made Mosaic to create proprietary browsers for them.

Sensing the opportunity, one of the Mosaic developers, Marc Andreessen, founded the company Mosaic Communications Corporation and created a new web browser named Mosaic Netscape. To resolve legal issues with NCSA, the company was renamed Netscape Communications Corporation and the browser Netscape Navigator. The Netscape browser improved on Mosaic's usability and reliability - as well as boasting the then-impressive feature of being able to display pages as they loaded.

Within a year or so, Netscape was the browser of choice for the Internet users. It had almost 90% user base covered. But then entered Microsoft's Internet Explorer 1.0, which was bundled free of cost with the Windows OS. This started, what is called, the first browser war. Both the companies, in a desire to outdo each other, started implementing more and more features to their browsers.

But still the web pages were static. Netscape decided to put user interactivity to the web pages.

They hired Brendan Eich to develop a scripting language that could interface with the server-side components. Tasked with this, Eich eventually decided that a loosely-typed scripting language suited the environment and audience, namely the few thousand web designers and developers who needed to be able to tie into page elements (such as forms, or frames, or images) without a bytecode compiler or knowledge of object-oriented software design.

Thus "LiveScript" was born. The name was chosen to reflect its dynamic nature. It was released with Navigator 2.0 but was quickly (before the end of Navigator 2.0 beta cycle) renamed to JavaScript.

Back after a long time...

Well... well... well... look who is back!!

Anyways, there are no readers to complain. So no apologies offered. Anyways, now to the main topic of today's post.

I am converting this blog to the chronicle of my quest for the Semantic Web.

The posts now will be related to Web 3.0 (Semantic Web) and various tools required to build a POC on my previous hypothesis, PHP, Javascript, MySql, CSS and Ajax.

Hope the world doesn't come apart before I complete my POC (What with the 2012 prophecy, you never know what's coming).

WEB 3.0 : The new frontier




I am an avid googler and swear by the Wikipedia, but a few days ago I was let down big time by the Google. I was to go to Ooty and was searching for suitable accommodation over there. Can you believe google gave me some 8470 results, out of which I read the first 20 but still couldn't get what I wanted.

Anyways, I decided that these kinda queries should be answered in person and no google should be given the authority to dictate the beds in which I sleep in. However, all the while I was driving to Ooty, one thought kept troubling me. Why can't someone dig all the wealth of information in the blogs and user reviews out there and provide me with a simple choice of no more than 10 hotels/homestays each in a different price range and with the maximum number of favorable reviews in their range?

This question is the ultimate frontier for Internet searching. Providing relevant search results have been the nightmare for all the search algorithm writers for over two decades now. Since the birth of the Internet, people have been collecting data on the cloud. How nice it would be to go through this huge amount of data and come back with the most relevant piece of information.

All this and some more spiked me to do some more googling on such searching and voila! I came up with a new jargon "WEB 3.0". Well, not actually new, I had heard a lot about semantic and vertical searching and have dirtied my hands trying web page - scraping but had never given a serious thought to this so called "natural language searching".

Now, lets take the most basic question : What is, or rather will be, WEB 3.0? Everybody has got his own opinion about what it will be, I also have mine. For me, WEB 3.0 will be a paradigm shift from what we know of the Internet as of today. I mean, 10 years down the line there won't be any website as we see today. What will be is a huge repository of data, which will be essentially user and community generated and we will be able to access the data in whatever format we like and we won't need a conventional computer to do that (anyways PCs will have shrunk to the size of a laser device mounted on our ears which can project images on any surface or play sounds from their ear buds). This data will be rendered in whatever format we like using our previous preferences and could be changed whenever we want to. Hmmm! Quite futuristic huh. Wait for another 10 years sweetheart.

But before we leap 10 years in the time warp, lets think if we can do anything about what we have here today. Maybe, maybe not. Lets break this huge insurmountable problem into somewhat smaller and manageable issues. As we know most of the searches in future will be like :

a. I want to go to a happy place.
b. I would like to read a sad, romantic story.
c. I would like to have delicious, Chinese, home cooked food.

Now all these questions today are answered by searching for keywords from the existing pages. However, we are talking about a search in which the search engine crawls all through the blogs and forums and other community sites and get the reactions of people for the various options possible and then show the results. Whew! That's a huge requirement in itself. Now let us break this into a further smaller problems.

The biggest issue here is how does the search engine understand the emotions portrayed by the web page? Most of the searches can be made easy if the pages can be ranked according to their EQ. To answer this let me ask another question: How do we gauge the emotional state of something we read? Simple! By looking at the keywords which our parents and teachers taught us to identify with certain emotions. Similarly if we keep a data base where in the search engine can refer to and find out the number of sad or happy keywords, then it can gauge the overall emotional state of the page.

So one problem solved. Similarly I will try to tackle different issues as and when I get time and in the end we will have the model of a basic WEB 3.0 search engine.

Oracle Buys Sun

Finally the beleaguered Sun Microsystems has got a suitor. Oracle Corp. has inked a deal to buy struggling server vendor Sun Microsystems for $9.50 per share, representing a 42.0% premium to Sun's closing price of $6.69 on Friday afternoon. This deal is worth $7.4 billion, or $5.6 billion net of Sun’s cash and debt.

What does it mean to us as a consumer? Now Oracle will own Sun's Solaris operating system, Java programming language, and servers and storage hardware systems businesses. Integrating its enterprise software products with now added server and Solaris OS it can provide a off-the-shelf ready-to-use product which you can buy, bring to your office, plugin to your network and voila! you are all set to start your applications on your server. These are exciting propositions. Oracle will be the only company with products like this, though IBM has its products on the same lines but the software products of Oracle have a greater demand.

Oracle is predicting that it will generate $2.0 billion in profits for Sun two years after signing the deal, which is expected to be sealed this summer. So it seems to be a very exciting buy after all. Looks like Larry Ellison is going to enhance his riches manifolds. He has taken a bold decision and only time will tell whether it is going to pay.

The Pirate Bay is pirated

Hmmmm... so a Napster happened to the Pirate Bay, though in a new country and by a new Judge. The defendants Frederik Neij, Gottfrid Svartholm Warg, Peter Sunde, and Carl Lundstrom were all sentenced to one year in jail each and a fine to the tune of $3,620,000 is to be paid by them to the prosecutors. The plaintiffs had demanded over 100 million Crowns as damages!!


Now, coming back to the issue of Piracy, I really don't take sides. I am of the notion that every man who works hard has a right to his creation and sell it at a price he deems to be right. But what about the obscene amounts these entertainment industry guys charge us. Why should I pay for the 20 millions some jackass named Brad Pitt charges these studios? Or for that matter the 20 Crores Akshay Kumar makes for one single movie. Am I as a spectator at fault if I can't shell out the Rs 250 my local friendly PVR charges me for one show? Why won't I go for a cheaper(or shud I say free) version on the Pirate bay, Mininova, Demonoid or Torrentbox? Only if the shows were at a reasonable rate like Rs 50 or perhaps lower than that I would obviously go and watch the movie on a 75 mm screen and not on my measly 17" PC Screen.

Lets take a hypothetical scenario : Movies are available for 1$ each for download from the net, using the same technologies. Will you then go for a pirated version? I won't!! The cost of any ones conscience is always more than a dollar. This is simple economics and I don't know why these guys don't understand this. Why bleed billions of Dollars to new technologies when you can actually earn millions from it? Instead of being steadfast on not lowering the prices, these studio bosses should acknowledge the advent of technologies and smartasses who use them (or abuse them depending on which side you are on) and understand the new demand and supply conditions. Gone are the days when you can only watch new movies in the theater or wait it to be at least 3-4 months older so that your favorite TV channel broadcasts it. With the faster than light broadband speeds the movies are passed on the net even before you can properly spell the cast's names.

The studio bosses should go in for differential pricing, the first three days should be at a premium of not more than $10. Then after the first weekend the prices should come down to $1. They don't have to maintain any servers or the huge bandwidth required by using the P2P distribution systems. All they need to do is to give a authenticated torrent tracker file for a dollar. And they also save money on printing discs and transportation. I sincerely believe this would be a highly profitable model.

And if they don't heed to it then no Pirate Bay can save them. Everyday the technology and browsing speed is increasing in scope. We will have hundreds of new sites giving away these so-called copyrighted stuff for free on the net and they can't sue everyone. Already there are sites which promise IP cloaking for P2P connections. And surprise surprise... Pirate Bay is still up and running. Looks like they are not ruffled. They have safely moved their servers to Netherlands!!

Someone needs to tell them to take a long hard thought at this age old saying "Lets make friends and not enemies".

How to break Win XP password?

This is one of the topics which I see frequently on any hacker community. So lets once and for all crack the SAM mystery.

What is SAM?

SAM (Security Accounts Manager) file stores all the user info and passwords of all the accounts of a computer using Windows NT family OS(Windows XP, Windows server 2003,etc.).So if you can somehow get this file you can get the passwords.

How can one find passwords from the SAM file?

There are three places where this file can be cracked from:-

i) From the original file
%systemroot%/system32/config
This file is locked to all users during the windows is running,so that you can't open it while you are working in windows. (Find out how you can use this file....Google dear friends).

ii) The system keeps a backup of this file in the
%systemroot%/repair/sam._
This file is available to all users at any time. So copy this file to any directory and crack the passwords using any good password cracker. I would tell you about one, not only coz its very popular but also that its free.(Find others yourselves the net has a gr8 many of them)

John the Ripper:- Its a dictionary cracker and will crack almost 80% of times you use it(unless the system admin has a knack in complicating things.)

iii) You can use PWDUMP to directly crack the passwords from the registry.pwdump uses .DLL injection in order to use the system account to view the password hashes stored in the registry.(Try to find out more about pwdump)

How to prevent people from cracking ur SAM file?

i) Try to avoid password which are dictionary words.

ii) Try to use special characters in ur password.

iii)Try to add non-printable ascii characters to your passwords.

Learning the OS's

Operating Systems form the heart and soul of computer systems. They are the set of programs which run the computer and help you process the data at such an incredible speed that it almost appears magical to everyone.

Now for anyone to be able to hack into any system will love to know the vulnerability of the OS running on the system. And so must you all(as u are aspiring hackers, right!). So lets start studying about OS in detail.

First and foremost study all there is to study about the basics of OS. The OS's may change but the concepts remain the same more or less.

Then install an earlier version of Windows(98 or 95). Now don't stare at me like that! I know that windows is BAD in terms of security, however it is a hackers heaven. You can try out all ur skills as a practice test here before venturing out. Play up with the OS, tweak the settings,registries and the codes(Did I say codes?). Try everything you can think of don't worry about crashing your system(Windows are meant to crash anyways). If you worry about failure then remember that failures are pillars of success (But try not to build only pillars without any hope to ever building the roof).This will give you thrill and the boost to move forward.

Then go for the *nixes. Try to install all variants of linuxs that you can lay hands on (Don't worry they can be downloaded from the net free of cost). Read all the MAN pages religiously. Read all that you can find about UNIX and LINUX (including their fascinating history). Try out hands-on on all that you can on these OS.

This completes the basics of hacking. Do these and remember these three E's

Explore Experiment and Enjoy

You will be a hacker in the true spirit of the word.

Don't worry there will be more on this blog by me. Just be a little more active and post more questions. It will help us both.

Learn Programming !!

So guys now that we have learnt how to use the net for our benefit,lets move on to the next level.

A hacker knows a number of programming languages.Those are his tools and believe me as there is a different tool for different situations, you will face situations where you find that you have to use a different programming language.

So here are a few languages which you have to have to know :-

i) You must know web designing languages including HTML,XHTML CSS,Javascript, vb script , mysql and php to learn to hack anything involving the internet. Also you may need to build a webpage where you will write about all your exploits. These languages are what you will need then.

ii) Thoroughly learn C or any of its variant(like C#, C++, etc.). By thoroughly I mean that you must be pretty good with programming big projects using this language, not just the "Hello World" stuff.

iii)Learn a scripting language like PERL(this is what I know, love and recommend but who cares for my advice,EH!!) or PYTHON. There are others but I don't know them and how can I recommend anything without first using it.

You will get a lot of good tutorials on the net for these languages. Just go to
Google and search for tutorials. A lot of them will be there choose one that suits your style.

Get source codes from the net(There aren't millions of them floating there, to be frank, but you can find some good ones there). Try to analyse them and tweak them for better output. Play with the codes a lot.

When you think you are ready you can get projects on the net and try to finish them to the end. You will love it when you finish them.

This is the end of todays topics.

Any doubts,questions,suggestions are are welcome.Feel free to express urselves guys.

Happy hacking!!!

How to become a Hacker?

So now that you know what is a hacker, don't you wanna know how to become one.

I will tell you this in a step by step method.

Everyday (or may be in two days,at the most,if I am busy) there will be an article as to how to become a hacker. The topics will be in an increasing order of intensity and interest. So I advice you to read them chronologically.

Now this topic concerns How to start?

This is the most frustrating part.You wanna learn something and there is no one to tell you how to do it.Don't worry I will give you an hint as to how to start. This is based on how I started and how most of "them" start.

Learn to use the Internet

You may have heard that the Net is a vast ocean of knowledge. But howcome you have never found it. Its because you have never ventured into deep waters. Try to do it and use a trustworthy search engine as your helmsman.

I would suggest these two search engines:-


Try to search every thing and look out for new and interesting things which you may have never looked at.

Searching the net for a particular piece of information is like searching for a needle in a hay stack. But don't worry. There are lots of tutorials on the net as to how to search properly. Google for it(this means search for it in Google, try to learn this expression as you will find it very frequently on the various newsgroups and forums).

Join Forums, Newsgroups and Mailing lists

You can learn best in a peer to peer arrangement. So join forums and then search out for people who you think are almost at your level of learning and start sharing with them.

And remember that you are new to the community so don't blabber anything which you are not confident about(Hackers have a very good memory and they don't forgive and forget mistakes).

Now, this is the end of this post. I'll back with some good post tomorrow. Till then happy hacking!!