Sora Innosia

Sora Innosia provides Free Softwares

Sora Innosia WebScrapper - Overview

What is Sora Innosia WebScrapper library?

WebScrapper is a library to quickly get data from webpage using only a single syntax! Yes, only a single syntax. Imagine developer have to scrap a website for certain information, developer has to write code such as

WebClient wc = new WebClient();
string googlecom = wc.DownloadString("");
string[] links = GetLinks(googlecom);            
string LINK1 = GetLink(links[5]);
string LINK2 = links[6];
string newData = wc.DownloadString(LINK1);

public string[] GetLinks(string content)
// ... Parsing of links

public string GetLink(string content)
string search = ",{t:5}); class=gbzt id=gb_5 href=\"";
int startIndex = content.IndexOf(search);
int endIndex = content.IndexOf("\"", startIndex + search.Length);
return content.Substring(startIndex + search.Length, endIndex - (startIndex + search.Length));

There are lots of lines which I even omit the method implementation to return list of links. By using WebScrapper library, you only need these lines

string syn = "SetResult('LINK1,LINK2', TagMatch(Filter(TagMatch(Download(''), '<a', ''), '5,6'), ',{t:5}); class=gbzt id=gb_5 href=\"', '\"'));Download(GetResult('LINK1'))";
WebScrapper.Scrapper scr = new WebScrapper.Scrapper();
string[] result = scr.Multiple(syn);

Simple? It is only 3 lines? Though the first lines look compact but the function itself is purely for retrieving web purposes, which including
1. Downloading from
2. Searching tag that matches '<a' and ''
3. Filter the result and return only item in index 5 and 6
4. Filter the result above that matches ',{t:5}); class="gbzt" id=gb_5 href=\"' and '\"' and it item 5 matches while item 6 does not match which will return empty string
5. Assign variable LINK1 equal to filtered item 5 and variable LINK2 equal to filtered item 6 (empty)
6. Download from variable Link1

The separation of concern meaning that the compact syntax serves only for one single purpose, which is Web Scrapping, which we have no control over the content since the content belong to other entity. We need a strong research and testing so that assumption is made that certain searches return certain result.

Many developer when doing Web Scrapping, assumes a lot of things, because there is no definite way that a website will stay as it is, the company behind the site, or entity behind the site, might do renovation, or works, upgrading or maintenance that cause the web changes. If we specifically write a .NET assembly such as DLL or EXE to get data based on our research, our DLL or EXE is easily outdated once the website doing changes, thus we have to analyze the website again and update our DLL or EXE code and doing recompilation and publish our code to our user or website. It is a tedious cycle that often happens.

By using WebScrapper, parsing of web is done using a single syntax which is a single string consisting of recursive statements. A string can be stored in database or configuration files, which makes it easy to modify without the need to recompile any code. When the target website changes, developer only needs to update the scrapping syntax and the scrapping works again!

Benefit of WebScrapper
1. Single Syntax in a string, thus can be stored in database or configuration files. Updating of Single Syntax is easy.
2. No need to compile the syntax, as it is being interpreted on the fly.
3. One instance of the Class uses one single WebClient control that maintain the Cookies state, thus downloading multiple page will keep the Cookies intact.
4. Support Regex
5. Built in string finder

And much more benefit when using WebScrapper instead of manual hard coding and compiling codes!
Download Now to test it!


20 Aug 2015 - We are Hiring for Moderator
Dear Users, We need anyone who is keen to work without pay for us to do any of below Upload missing...
19 Aug 2015 - Request Reupload and New Upload System Enhancement
Dear Users, Previously we will only able to let you reupload file to your File Host account when the...
7 Aug 2015 - A Lot Has Been Fixed
Dear Users, A lot has been fixed after Odin notify me with the error. You might encounter below...
5 Aug 2015 - Has Been Fixed
Dear Users, You may have noticed recently mega link is missing and this has been caused by...
1 August 2015 - Mass Upload Slow Down Encode
Dear Users, Please expect of 25% slower download since server is currently busy uploading files to...