2014年6月1日 星期日

[NodeJS] 透過 Wget + HTTP Proxy Server 架構進行 Content 分析

前陣子研究抓取資料時,發現 web crawler 搭配 proxy 時,可以在 proxy 這層進行資料的處理、分析,就一直思考到底要用 python 還是 node.js 測試,最後,實在是 node.js 方便許多,就...衝啦。

在進行之前,小提一下 Proxy Server 本身就有 Internet Content Adaptation Protocol (ICAP) 機制,其實只需要架一個 Squid 在寫一些 ICAP 即可達到類似的效果,有興趣可參考:[Python] 使用 PyICAP 淺玩 ICAP 與 Squid (以 Response Modification / RESPMOD 為例) @ Ubuntu 12.04。

對於沒開發過 node.js 的我,就先參考 node-http-proxy 的目錄結構建立起來一個 open source:node-content-filter-proxy。

這兩天抽空把玩的心得:
  • 支援 HTTP Requests
  • 尚未處理 HTTPS Requests

抽取 <a href="...">...</a> 用法:


$ cd node-content-filter-proxy/examples && npm install
$ clear;node extract-hmtl-a-href.js
1 Jun 10:20:22 - Service Running at 3128 Port


透過 wget proxy usage:

$ http_proxy="http://localhost:3128" https_proxy="http://localhost:3128" wget -qO -  http://www.google.com/
<a href="http://www.google.com.sg/imghp?hl=en&tab=wi">Images<a/>
<a href="http://maps.google.com.sg/maps?hl=en&tab=wl">Maps<a/>
<a href="https://play.google.com/?hl=en&tab=w8">Play<a/>
...


簡單的說,想要透過既有的 crawler (wget) 去下載資料,但是,資料中我只對 hyperlink 感到興趣,所以透過 proxy server 更新成只回傳 hyperlink 就好,維持 hyperlink 結構可以讓 wget --recursive 抓資料。

如此一來,原先 crawer = wget 架構,轉形成 crawler = wget + proxy server + content analysis 模式,可以專心擺重在內容分析,不必從頭刻一隻類似 wget 的程式。

最後一提,在 extract-hmtl-a-href.js 範例中,採用的不是透過 regular expression 去分析 link ,而是接近一隻 Javascript Rendering Engine (cheerio) ,所以未來可以做的事更多了 :) 至於效率部分倒還好,畢竟 crawler 抓太快會被 ban 掉,此外,分析 content 時,也能設計複雜更高的 seed list (回吐  hyperlink 時),能避開 DOS 現象(例如限定某個 domain 、網站一天只抓2000筆等)。

2014年5月29日 星期四

[Linux] 讓 git checkout/diff 略過 file permission 屬性

最近把 git 當作 deployment 過程的一環,接著就會碰到一個問題,有時因為需求要調整 file permission ,此時就會造成 git checkout 失敗,要求要先 merge 在行動。雖然 file permission 容易出現在 Windows / Linux ,但在 Linux 環境用 git 當作 deployment 時,也常有身份、權限的改變而造成一樣的影響。

解法:

$ git config core.fileMode false

或

$ git -c core.fileMode=false diff
$ git -c core.fileMode=false checkout

2014年5月25日 星期日

[NodeJS] 使用 WebShot 進行網頁截圖、顯示正確的中文(CJK)等編碼 @ Ubuntu 14.04 Server



把玩一下 node.js ,發現有個套件不錯叫做 webshot,非常簡易地就可以把網站截取下來。然而,在 Ubuntu 14.04 server 版上運行時,發現無法正常顯示中文字,需要額外處理,就筆記一下:

安裝 Node.js:

$ sudo apt-get install nodejs npm
$ sudo ln -s /usr/bin/nodejs  /usr/bin/node


使用 webshot:

$ mkdir ~/webshot
$ cd ~/webshot
$ npm install webshot
$ vim test.js
var webshot = require('webshot');
webshot('tw.yahoo.com', 'yahoo.png', function(err) {
if(err)
       console.log(err);
} );
$ nodejs test.js


然而,無法顯示中文字,簡言之就是缺字型,安裝一下即可:

$ sudo apt-get install xfonts-wqy

成果一切正常:

2014年5月23日 星期五

使用 Wget 進行簡易的 Web Crawling

$ wget --recursive --no-clobber --html-extension --convert-links --no-parent --wait 5 --domains example.com www.example.com

如此一來,會在當前目錄建立 www.example.com 目錄,將依照網站目錄結構存儲起來。

使用 lnyx 抽取 Web Page 所有的 Links

沒想到 lynx 真好用 XD 比用 Wget 好的地方是不用去重組一些 relative link,那缺點就會是要避開同一個 page 的 anchor (<a href="#me></a>) 用法,但是...有些 Javascript 的或是其他 MVC 架構的網站,仍會用 anchor 來取資料。

$ lynx -dump -listonly https://tw.yahoo.com | grep -o '^\s\{1,\}[0-9]\{1,\}..*$' | sed -e 's/^[[:space:]]\{1,\}[0-9]\{1,\}\.[[:space:]]//g' | uniq