프로그래머 아닌 분들께는 해당없습니다.
Url을 주면 title, description, image를 읽어내 주는 코드를 만들었습니다.
백문이 불여일견, 백견이 불여일타.
playframework scala로 되어 있고, 프로젝트는 https://github.com/Aha00a/crawler 에 있습니다.
중심로직은 https://github.com/…/c…/blob/master/app/logics/Crawler.scala 입니다.
기본적으로 html title과 meta og를 파싱하며, 네이x 인물정보의 경우 수동처리하여 정상적인 결과가 나오도록 하였습니다.
필요하신분 있으시면 참고하세요.
Similar pages by cosine similarity. Words after page name are term frequency.