"""Parse only explicit component evidence; never infer actual members from a preview.""" import re from datetime import datetime from uuid import uuid4 from fastapi import HTTPException from sqlalchemy import select from ..alphas import sanitize from ..backtests.contracts import fingerprint from ..models import SuperSelectionSnapshot from ..research.serialization import encode_snapshot def parse_components(raw): """Return normalized rows and completeness; count/duplicate/next ambiguity stays unknown.""" warnings = [] if isinstance(raw, dict): supplied = raw.get("warnings", []) warnings.extend(supplied if isinstance(supplied, list) else [supplied]) rows = raw.get("results", raw.get("alphas", raw.get("components"))) total = raw.get("count", raw.get("total")) complete_hint = raw.get("complete") is True next_page = raw.get("next") else: rows, total, complete_hint, next_page = raw, None, False, None invalid_total = total is not None and (type(total) is not int or total < 0) total = total if type(total) is int and total >= 0 else None valid_shape = isinstance(rows, list) items, seen, malformed = [], set(), False for row in rows if valid_shape else []: entry = {"id": row} if isinstance(row, str) else row if not isinstance(entry, dict): malformed = True continue alpha_id = entry.get("id", entry.get("alpha", entry.get("alphaId"))) if not isinstance(alpha_id, str) or not re.fullmatch(r"[A-Za-z0-9_-]{1,100}", alpha_id) or alpha_id in seen: malformed = True continue seen.add(alpha_id) items.append({**sanitize(entry), "id": alpha_id}) complete = valid_shape and not malformed and not invalid_total and not next_page and ( (total is not None and total == len(items)) or (total is None and complete_hint)) if not complete: warnings.append("组件列表未核实完整性;不生成完整组件指纹,不用于同池结论") return {"components": items, "total": total, "complete": complete, "component_hash": fingerprint({"alpha_ids": sorted(seen)}) if complete else None, "warnings": sanitize(warnings)} def snapshot_output(row, limit=25, offset=0, q=""): items = [item for item in row.components if not q or q.lower() in str(item).lower()] return encode_snapshot({"snapshot_id": row.id, "job_id": row.job_id, "item_id": row.item_id, "source": row.source, "request": row.request, "request_hash": row.request_hash, "component_hash": row.component_hash, "complete": row.complete, "reported_total": row.total, "observed_at": row.observed_at, "warnings": row.warnings, "status": "available" if row.complete else "unknown", "total": len(items), "limit": limit, "offset": offset, "has_more": offset + limit < len(items), "items": items[offset:offset + limit]}) async def read_selection(db, args): query = select(SuperSelectionSnapshot) query = query.where(SuperSelectionSnapshot.id == args.snapshot_id) if args.snapshot_id else query.where( SuperSelectionSnapshot.job_id == args.job_id) row = await db.scalar(query) if not row: if args.job_id: from ..models import Job job = await db.get(Job, args.job_id) if not job or job.kind != "super_selection_preview": raise HTTPException(404, "组件预览任务不存在") return {"status": job.status, "snapshot_id": None, "job_id": job.id, "items": [], "total": 0, "complete": False, "error": job.error, "observed_at": None} raise HTTPException(404, "组件快照不存在") return snapshot_output(row, args.limit, args.offset, args.q) async def save_actual_components(db, item, detail, observed_at): raw = detail.get("components", detail.get("selectedAlphas")) if raw is None and isinstance(detail.get("selection"), dict): selection = detail["selection"] if isinstance(selection.get("alphas"), list): raw = {"alphas": selection["alphas"], "count": selection.get("count")} request = {"type": "SUPER", "selection": item.selection, "combo": item.combo, "settings": item.settings} parsed = parse_components(raw) db.add(SuperSelectionSnapshot(id=str(uuid4()), item_id=item.id, source="actual", request=request, request_hash=fingerprint(request), raw=sanitize(raw) if isinstance(raw, (dict, list)) else {}, observed_at=datetime.fromisoformat(observed_at), **parsed)) async def actual_components(db, item_id, limit=25, offset=0): row = await db.scalar(select(SuperSelectionSnapshot).where(SuperSelectionSnapshot.item_id == item_id)) return snapshot_output(row, limit, offset) if row else { "status": "unknown", "complete": False, "source": "actual", "items": [], "total": 0, "component_hash": None, "observed_at": None, "warnings": ["平台实际组件尚未核实"]}