Quay lại danh sách
Xây dựng hệ thống đánh giá proxy tự động để tối ưu scraping lớn

Xây dựng hệ thống đánh giá proxy tự động để tối ưu scraping lớn

30 tháng 9, 2026

Tại sao cần hệ thống đánh giá proxy tự động?

Khi scraping quy mô lớn, pool proxy thường chứa hàng nghìn IP. Không phải IP nào cũng ổn định: một số bị chặn, một số chậm, một số bị đánh dấu spam. Việc chọn ngẫu nhiên hoặc theo round‑robin dẫn đến tỷ lệ thất bại cao, lãng phí băng thông và thời gian xử lý lỗi. Một hệ thống health‑score (điểm sức khỏe) tự động giúp:

  • Loại bỏ IP kém chất lượng trước khi request thật.
  • Ưu tiên IP nhanh, sạch cho các tác vụ nhạy cảm (đăng nhập, thanh toán).
  • Giảm chi phí bằng cách tận dụng tối đa proxy tốt, giảm số lượng proxy cần thuê.

Các chỉ số cốt lõi cho health‑score

Chỉ số Ý nghĩa Cách đo lường Trọng số gợi ý
Latency (ms) Thời gian phản hồi trung bình đến target GET /health hoặc request thật đến trang mục tiêu, lấy response.elapsed 0.35
Success Rate (%) Tỷ lệ request trả về 2xx/3xx trong cửa sổ quan sát Đếm thành công / tổng request trong 5‑10 phút gần nhất 0.40
IP Reputation Điểm uy tín từ các dịch vụ bên thứ ba (AbuseIPDB, IPQualityScore) Gọi API reputation, chuẩn hóa về 0‑100 0.15
Geo‑match Độ phù hợp vị trí địa lý so với target So sánh country_code của proxy với quốc gia mục tiêu 0.10

Lưu ý: Trọng số có thể điều chỉnh theo nghiệp vụ. Ví dụ scraping giá cả cần latency thấp → tăng trọng số latency.

Kiến trúc tổng quan

+----------------+      +-------------------+      +------------------+
|  Proxy Pool    | ---> |  Metric Collector | ---> |  Scoring Engine  |
|  (Redis/DB)    |      |  (cron / worker)  |      |  (Python/Node)   |
+----------------+      +-------------------+      +------------------+
                                                          |
                                                          v
                                               +----------------------+
                                               |  Proxy Selector API  |
                                               |  (FastAPI / Express) |
                                               +----------------------+
                                                          |
                                                          v
                                               +----------------------+
                                               |  Scraper Workers     |
                                               |  (requests/axios)    |
                                               +----------------------+
  • Proxy Pool lưu trữ metadata: ip, port, username, password, country, type (residential/datacenter/mobile).
  • Metric Collector chạy định kỳ (cron mỗi 2‑5 phút) gửi request kiểm tra tới một endpoint nhẹ (ví dụ https://httpbin.org/ip) qua từng proxy, ghi lại latency, status code, error.
  • Scoring Engine đọc metrics gần nhất, tính điểm theo công thức weighted sum, cập nhật health_score vào pool.
  • Proxy Selector API cung cấp endpoint GET /best-proxy?country=US&min_score=70 trả về proxy tốt nhất phù hợp.
  • Scraper Workers gọi Selector API trước mỗi batch request.

Triển khai Metric Collector (Python)

# collector.py
import asyncio
import aiohttp
import time
import json
import redis
from dataclasses import dataclass
from typing import List

REDIS = redis.Redis(decode_responses=True)
TARGET = "https://httpbin.org/ip"          # endpoint kiểm tra nhẹ
CONCURRENCY = 50
TIMEOUT = aiohttp.ClientTimeout(total=10)

@dataclass
class Proxy:
    ip: str
    port: int
    user: str
    pwd: str
    country: str
    ptype: str

async def fetch(session: aiohttp.ClientSession, proxy: Proxy) -> dict:
    proxy_url = f"http://{proxy.user}:{proxy.pwd}@{proxy.ip}:{proxy.port}"
    start = time.perf_counter()
    try:
        async with session.get(TARGET, proxy=proxy_url, timeout=TIMEOUT) as resp:
            latency = (time.perf_counter() - start) * 1000  # ms
            ok = 200 <= resp.status < 400
            return {"ip": proxy.ip, "latency": latency, "success": ok}
    except Exception as e:
        return {"ip": proxy.ip, "latency": None, "success": False, "error": str(e)}

async def run():
    # Lấy danh sách proxy từ Redis hash "proxies"
    raw = REDIS.hgetall("proxies")
    proxies: List[Proxy] = []
    for k, v in raw.items():
        data = json.loads(v)
        proxies.append(Proxy(**data))

    connector = aiohttp.TCPConnector(limit=CONCURRENCY)
    async with aiohttp.ClientSession(connector=connector) as session:
        tasks = [fetch(session, p) for p in proxies]
        results = await asyncio.gather(*tasks)

    # Ghi metric vào sorted set theo IP, giữ 100 mẫu gần nhất
    for r in results:
        ip = r["ip"]
        now = int(time.time())
        member = json.dumps({"ts": now, "latency": r["latency"], "success": r["success"]})
        REDIS.zadd(f"metrics:{ip}", {member: now})
        # Trim cũ
        REDIS.zremrangebyscore(f"metrics:{ip}", 0, now - 3600)  # giữ 1h

if __name__ == "__main__":
    asyncio.run(run())
  • Chạy script này qua cron */5 * * * * /usr/bin/python3 collector.py.
  • Sử dụng aiohttp bất đồng bộ để kiểm tra hàng trăm proxy cùng lúc.
  • Metrics lưu trong Redis sorted set với timestamp làm score → dễ dàng lấy N mẫu gần nhất.

Scoring Engine – tính health_score

# scoring.py
import redis
import json
import statistics

REDIS = redis.Redis(decode_responses=True)
WEIGHTS = {
    "latency": 0.35,
    "success": 0.40,
    "reputation": 0.15,
    "geo": 0.10,
}

REPUTATION_CACHE = {}

def get_reputation(ip: str) -> float:
    """Trả về 0‑100, cache 6h."""
    if ip in REPUTATION_CACHE:
        return REPUTATION_CACHE[ip]
    # Giả lập gọi API AbuseIPDB
    # resp = requests.get(f"https://api.abuseipdb.com/api/v2/check?ipAddress={ip}", headers={...})
    # score = 100 - resp.json()['data']['abuseConfidenceScore']
    score = 85.0  # placeholder
    REPUTATION_CACHE[ip] = score
    return score

def compute_score(ip: str, target_country: str = None) -> float:
    # Lấy 30 mẫu gần nhất
    members = REDIS.zrevrange(f"metrics:{ip}", 0, 29)
    if not members:
        return 0.0
    latencies = []
    successes = []
    for m in members:
        data = json.loads(m)
        if data["latency"] is not None:
            latencies.append(data["latency"])
        successes.append(1 if data["success"] else 0)
    avg_latency = statistics.mean(latencies) if latencies else 9999
    success_rate = sum(successes) / len(successes) * 100
    # Chuẩn hóa latency: 0‑100 (càng thấp càng tốt)
    latency_score = max(0, 100 - (avg_latency / 200) * 100)  # 200ms → 0 điểm
    rep_score = get_reputation(ip)
    geo_score = 100 if (target_country is None or REDIS.hget(f"proxy:{ip}", "country") == target_country) else 0

    score = (
        WEIGHTS["latency"] * latency_score +
        WEIGHTS["success"] * success_rate +
        WEIGHTS["reputation"] * rep_score +
        WEIGHTS["geo"] * geo_score
    )
    return round(score, 2)

def update_all_scores(target_country: str = None):
    for key in REDIS.scan_iter("metrics:*"):
        ip = key.split(":", 1)[1]
        sc = compute_score(ip, target_country)
        REDIS.hset(f"proxy:{ip}", "health_score", sc)

if __name__ == "__main__":
    update_all_scores()
  • Hàm compute_score chuẩn hóa latency về thang 0‑100, kết hợp với success‑rate, reputation, geo‑match.
  • Chạy scoring mỗi 5‑10 phút sau khi collector xong.

Proxy Selector API (FastAPI)

# selector.py
from fastapi import FastAPI, Query
import redis
import json

app = FastAPI()
REDIS = redis.Redis(decode_responses=True)

@app.get("/best-proxy")
def best_proxy(
    country: str = Query(None),
    min_score: float = Query(60),
    limit: int = Query(1),
):
    # Lấy tất cả proxy có health_score >= min_score
    candidates = []
    for key in REDIS.scan_iter("proxy:*"):
        data = REDIS.hgetall(key)
        if not data:
            continue
        sc = float(data.get("health_score", 0))
        if sc < min_score:
            continue
        if country and data.get("country") != country.upper():
            continue
        candidates.append((sc, data))
    # Sắp xếp giảm dần score
    candidates.sort(key=lambda x: x[0], reverse=True)
    top = [c[1] for c in candidates[:limit]]
    return {"proxies": top}
  • Deploy với uvicorn selector:app --host 0.0.0.0 --port 8000.
  • Scraper chỉ cần GET http://selector:8000/best-proxy?country=VN&min_score=70.

Tích hợp vào Scraper Worker (Python requests)

# scraper.py
import requests
import os

SELECTOR = os.getenv("SELECTOR_URL", "http://localhost:8000/best-proxy")

def get_proxy(country: str = "US", min_score: float = 70):
    r = requests.get(SELECTOR, params={"country": country, "min_score": min_score}, timeout=5)
    r.raise_for_status()
    proxies = r.json()["proxies"]
    if not proxies:
        raise RuntimeError("No healthy proxy")
    p = proxies[0]
    return {
        "http": f"http://{p['username']}:{p['password']}@{p['ip']}:{p['port']}",
        "https": f"http://{p['username']}:{p['password']}@{p['ip']}:{p['port']}",
    }

def scrape(url: str, country: str = "US"):
    proxies = get_proxy(country)
    resp = requests.get(url, proxies=proxies, timeout=15)
    resp.raise_for_status()
    return resp.text

if __name__ == "__main__":
    html = scrape("https://example.com/products", country="VN")
    print(html[:200])
  • Worker tự động lấy proxy khỏe nhất trước mỗi request.
  • Có thể mở rộng: retry với proxy khác nếu bị 403/429.

Mở rộng: Node.js version cho team frontend

// selector-client.js
const axios = require('axios');
const SELECTOR = process.env.SELECTOR_URL || 'http://localhost:8000/best-proxy';

async function getBestProxy(country = 'US', minScore = 70) {
  const { data } = await axios.get(SELECTOR, { params: { country, min_score: minScore } });
  if (!data.proxies.length) throw new Error('No healthy proxy');
  const p = data.proxies[0];
  return `http://${p.username}:${p.password}@${p.ip}:${p.port}`;
}

module.exports = { getBestProxy };
// scraper.js
const axios = require('axios');
const { getBestProxy } = require('./selector-client');

async function scrape(url, country = 'US') {
  const proxy = await getBestProxy(country);
  const resp = await axios.get(url, {
    proxy: { host: new URL(proxy).hostname, port: Number(new URL(proxy).port), auth: { username: new URL(proxy).username, password: new URL(proxy).password } },
    timeout: 15000,
  });
  return resp.data;
}

scrape('https://example.com/products', 'VN').then(console.log).catch(console.error);

Giám sát & Cảnh báo

  • Prometheus exporter đọc health_score từ Redis, đẩy metric proxy_health_score{ip="...",country="..."}.
  • Tạo alert: proxy_health_score < 50 for 5m → Slack/Email.
  • Dashboard Grafana hiển thị phân phối score theo quốc gia, loại proxy.

Best practice & Mẹo thực chiến

  1. Warm‑up proxy mới: trước khi cho vào pool, chạy collector 10‑15 phút để có metric ban đầu.
  2. Sticky session cho login: khi cần duy trì cookie, dùng endpoint /sticky-proxy?session_id=abc trả về cùng IP trong 10‑15 phút (cấu hình ở RoProxy).
  3. Rate‑limit per IP: kết hợp token‑bucket ở scraper để không gửi quá nhiều request/giây trên một IP.
  4. Fallback chain: nếu selector không trả proxy (pool cạn), tự động mở rộng min_score xuống 40 hoặc bật datacenter proxy làm dự phòng.
  5. Log correlation: ghi proxy_ip, health_score, latency vào log request để debug sau này.

Kết luận

Xây dựng hệ thống health‑score proxy biến pool proxy từ "hộp đen" thành tài nguyên có thể quan sát, đo lường và tối ưu hóa. Quy trình:

  1. Collector chạy liên tục → metric thực tế.
  2. Scoring Engine tính điểm đa chiều.
  3. Selector API cung cấp proxy tốt nhất theo ngữ cảnh (quốc gia, ngưỡng điểm).
  4. Scraper workers tiêu dùng API, tự động retry/fallback.
  5. Monitoring & alert đảm bảo chất lượng pool không giảm sút.

Với ~200 dòng code Python + FastAPI + Redis, bạn có một pipeline production‑ready, mở rộng được cho hàng chục nghìn proxy và dễ dàng tích hợp vào bất kỳ stack scraping nào (Python, Node, Go, Rust). Hãy bắt đầu bằng một collector nhỏ, quan sát metric, sau đó bật scoring và selector – hiệu suất scraping sẽ tăng rõ rệt trong tuần đầu tiên.